Selective inter-component transformation (ICT) for image and video coding
By selectively applying inter-component joint coding/decoding methods with explicit and implicit signaling, the proposed solution addresses the limitations of existing techniques, enhancing coding efficiency and flexibility in image and video coding.
Patent Information
- Application Number
- JP2025030302
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-12
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing inter-component prediction techniques for image and video coding, such as CCLM and JCC, face limitations in computational complexity and flexibility, particularly when applied between chroma channels or for RGB-coded inputs with significant chromatic aberrations.
The proposed solution involves selectively applying inter-component joint coding/decoding methods using explicit signaling and non-binary indices, allowing for the use of multiple inter-component transforms (ICT) and implicit signaling through existing coded block flags, thereby enhancing flexibility and reducing computational complexity.
This approach improves coding efficiency by allowing for the selective application of different inter-component transforms and implicit signaling, which reduces computational overhead and enhances adaptability to various color spaces and image content.
Smart Images

Figure 2025081704000001_ABST
Abstract
Description
[Technical field]
[0001] The following description of the figures begins with the presentation of an encoder and decoder of a block-based predictive codec for coding pictures of a video to form an example of a coding framework in which embodiments of the present invention can be incorporated. The respective encoders and decoders are described with reference to Figures 1 to 3. Below, a description of embodiments of the inventive concepts is presented together with a description of how such concepts can be incorporated into the encoders and decoders of Figures 1 and 2, respectively, although the embodiments described in subsequent Figures 4 onwards may also be used to form encoders and decoders that do not operate according to the coding framework underlying the encoders and decoders of Figures 1 and 2.
[0002] Elements that are equal or similar or have equal or similar functionality are designated in the following description with equal or similar reference numbers even if they occur in different figures.
[0003] In the following description, a number of details are set forth to provide a more complete description of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present invention. Furthermore, the features of different embodiments described below can be combined with each other unless otherwise specified.
[0004] FIG. 1 shows an apparatus for predictively coding a picture 12 into a data stream 14, exemplarily using transform-based residual coding. The apparatus or encoder is indicated using the reference number 10. FIG. 2 shows a corresponding decoder 20, i.e. an apparatus 20 configured to predictively decode a picture 12' from the data stream 14, also using transform-based residual decoding, the apostrophe being used to indicate that the picture 12' reconstructed by the decoder 20 deviates from the picture 12 originally coded by the apparatus 10 in terms of coding losses introduced by quantization of the predictive residual signal. Although FIG. 1 and FIG. 2 exemplarily use transform-based predictive residual coding, the embodiments of the present application are not limited to this type of predictive residual coding. This also applies to other details described with respect to FIG. 1 and FIG. 2, as outlined below.
[0005] The encoder 10 is configured to perform a spatio-spectral transform of the prediction residual signal and to encode the prediction residual signal thus obtained into a data stream 14. Similarly, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and to perform a spectro-spatial transform of the prediction residual signal thus obtained.
[0006] Internally, the encoder 10 may comprise a prediction residual signal former 22 which generates a prediction residual 24 in order to measure the deviation of a prediction signal 26 from an original signal, i.e. the picture 12. The prediction residual signal former 22 may for example be a subtractor which subtracts the prediction signal from the original signal, i.e. from the picture 12. The encoder 10 then further comprises a transformer 28 which performs a spatial-spectral transformation of the prediction residual signal 24 in order to obtain a spectral domain prediction residual signal 24' which is quantized by a quantizer 32 also included in the encoder 10. The prediction residual signal 24'' thus quantized is coded into the bitstream 14. For this purpose, the encoder 10 may optionally comprise an entropy coder 34 which entropy codes the prediction residual signal which is transformed and quantized into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 on the basis of the prediction residual signal 24'' which is coded into the data stream 14 and decodable therefrom. For this purpose, the prediction stage 36 may internally comprise, as shown in FIG. 1, an inverse quantizer 38 for inverse quantizing the prediction residual signal 24″ to obtain a spectral domain prediction residual signal 24″ corresponding to the signal 24′ but without the quantization losses, and an inverse transformer 40 for inversely transforming, i.e. spectral-spatial transforming, the latter prediction residual signal 24″ to obtain a prediction residual signal 24′″ corresponding to the original prediction residual signal 24 but without the quantization losses. A combiner 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24″″, for example by addition, to obtain a reconstructed signal 46, i.e. a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12′. A prediction module 44 of the prediction stage 36 then generates a prediction signal 26 based on the signal 46, for example using a spatial prediction, i.e. an intra-picture prediction, and / or a temporal prediction, i.e. an inter-picture prediction.
[0007] Similarly, as shown in Figure 2, the decoder 20 may be internally composed of components corresponding to the prediction stage 36, interconnected in a manner corresponding to the prediction stage. In particular, an entropy decoder 50 of the decoder 20 may entropy decode a quantized spectral domain prediction residual signal 24" from the data stream, with an inverse quantizer 52, an inverse transformer 54, a combiner 56, and a prediction module 58 interconnected and cooperating in the manner described above with respect to the modules of the prediction stage 36 to recover a reconstructed signal based on the prediction residual signal 24", such that an output of the combiner 56 provides a reconstructed signal, i.e., picture 12', as shown in Figure 2.
[0008] Although not specifically described above, it is readily apparent that the encoder 10 can set several coding parameters, including, for example, prediction modes, motion parameters, etc., according to several optimization schemes, such as, for example, methods that optimize several rate- and distortion-related criteria, i.e., coding cost. For example, the encoder 10 and the decoder 20 and corresponding modules 44, 58 can each support different prediction modes, such as intra-coded and inter-coded modes. The granularity at which the encoder and the decoder switch between these prediction mode types may correspond to a subdivision of the pictures 12 and 12', respectively, into coded segments or coded blocks. In units of these coded segments, for example, the picture may be subdivided into intra-coded and inter-coded blocks. The intra-coded blocks are predicted based on the already coded / decoded neighborhood of the respective blocks' space, as outlined in more detail below. Several intra-coding modes may exist and be selected for each intra-coded segment, including directional or angular intra-coding modes, whereby each segment may be filled by extrapolating neighboring sample values along a particular direction specific to each directional intra-coding mode to each intra-coded segment. The intra-coding modes may also include one or more further modes, such as, for example, a DC coding mode, whereby the prediction of each intra-coded block assigns a DC value to all samples in the respective intra-coded segment, and / or a planar intra-coding mode, along which the prediction of each block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of each intra-coded block with a planar driving slope and offset defined by the two-dimensional linear function based on neighboring samples. In comparison, inter-coded blocks may, for example, be predicted temporally.For inter-coded blocks, motion vectors can be signaled in the data stream, which indicate the spatial displacement of parts of previously coded pictures of the video to which picture 12 belongs, and the previously coded / decoded pictures are sampled to obtain a prediction signal for the respective inter-coded block. This means that in addition to the residual signal coding included in data stream 14, such as entropy-coded transform coefficient levels representing the quantized spectral domain prediction residual signal 24'', data stream 14 can encode some prediction parameters of the blocks, such as coding mode parameters for assigning coding modes to the various blocks, motion parameters of the inter-coded segments, and optional further parameters, such as parameters for controlling and signaling the subdivision of pictures 12 and 12' into respective segments. Decoder 20 uses these parameters to subdivide the picture in the same way as the encoder did, assign the same prediction modes to the segments, and perform the same predictions resulting in the same prediction signals.
[0009] FIG. 3 shows the relationship between the reconstructed signal, i.e. the reconstructed picture 12′, on the one hand, and the combination of a prediction residual signal 24′″ and a prediction signal 26 signaled in the data stream 14, on the other hand. As already mentioned above, the combination may be additive. The prediction signal 26 is shown in FIG. 3 as a subdivision of the picture area into intra-coded blocks, exemplarily shown using line shading, and inter-coded blocks, exemplarily shown without line shading. The subdivision may be any subdivision, such as a regular subdivision of the picture area into rows and columns of square or non-square blocks, or a multi-tree subdivision of the picture 12 from a tree root block into several leaf blocks of different sizes, such as a quad-tree subdivision, a mixture of which is shown in FIG. 3, where the picture area is first subdivided into rows and columns of a tree root block, which are then further subdivided into one or several leaf blocks according to a recursive multi-tree subdivision.
[0010] Again, data stream 14 may have an intra-coding mode coded for intra-coded blocks 80, which assigns one of several supported intra-coding modes to each intra-coded block 80. For inter-coded blocks 82, data stream 14 may have one or more motion parameters coded. Generally speaking, inter-coded blocks 82 are not limited to being temporally coded. Alternatively, inter-coded blocks 82 may be any block predicted from a previously coded portion beyond current picture 12 itself, such as a previously coded picture of the video to which picture 12 belongs, or a picture of another view or hierarchically lower layer if the encoder and decoder are scalable encoder and decoder, respectively.
[0011] The prediction residual signal 24'''' in FIG. 3 is also shown as a subdivision of the picture domain into blocks 84. These blocks may be called transform blocks to distinguish them from the coded blocks 80 and 82. In fact, FIG. 3 shows that the encoder 10 and the decoder 20 may use two different subdivisions of the picture 12 and the picture 12' into blocks, one into coded blocks 80 and 82, and the other into transform blocks 84. Although both subdivisions may be the same, i.e. each coded block 80 and 82 may simultaneously form a transform block 84, FIG. 3 shows the case where the subdivision into transform blocks 84 forms an extension of the subdivision into the coded blocks 80, 82, for example, such that any boundary between the two blocks 80 and 82 covers the boundary between the two blocks 84, or each block 80, 82 coincides with one of the transform blocks 84 or coincides with a cluster of transform blocks 84. However, these subdivisions may also be determined or selected independently of each other, such that the transformed block 84 may alternatively cross a block boundary between the blocks 80, 82. Thus, as far as the subdivision into transformed block 84 is concerned, similar statements are true as those presented with regard to the subdivision into blocks 80, 82, i.e., the block 84 may be the result of a regular subdivision of the picture area into blocks (with or without arrangement into rows and columns), a recursive multi-tree subdivision of the picture area, or a combination thereof, or any other kind of blocking. It should be noted that the blocks 80, 82, and 84 are not limited to being square, rectangular, or any other shape.
[0012] 3 further illustrates that the combination of the prediction signal 26 and the prediction residual signal 24'''' directly results in the reconstructed signal 12'. However, it should be noted that according to alternative embodiments, multiple prediction signals 26 can be combined with the prediction residual signal 24'''' into the picture 12'.
[0013] In Fig. 3, the transform blocks 84 are assumed to have the following importance: the transformer 28 and the inverse transformer 54 perform transformations in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow to skip the transformation for some of the transform blocks 84, so that the prediction residual signal is directly coded in the spatial domain. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured such that they support several transformations. For example, the transformations supported by the encoder 10 and the decoder 20 may include: DCT-II (or DCT-III), where DCT stands for discrete cosine transform DST-IV, where DST stands for discrete sine transform DCT-IV DST-VII type Identity Transformation (IT)
[0014] Of course, the transformer 28 supports all of the forward transform versions of these transforms, but the decoder 20 or inverse transformer 54 supports the corresponding reverse or inverse transform versions: · Reverse DCT-II (or Reverse DCT-III) Reverse DST-IV ·Reverse DCT-IV Reverse DST-VII Identity Transformation (IT)
[0015] The following description provides further details regarding which transforms may be supported by the encoder 10 and the decoder 20. Note that in any case, the set of supported transforms may include only one transform, such as one spectral-to-spatial or spatial-to-spectral transform. [Brief description of the drawings]
[0016] As already outlined above, Figures 1-3 are presented as examples in which the inventive concepts further described below can be implemented to form specific examples of encoders and decoders according to the present application. To that extent, the encoders and decoders of Figures 1 and 2, respectively, can represent possible implementations of the encoders and decoders described later in this specification. However, Figures 1 and 2 are merely examples. However, an encoder according to an embodiment of the present application can perform block-based encoding of the picture 12 using concepts outlined in more detail below that differs from the encoder of Figure 1, for example being a still picture encoder rather than a video encoder, not supporting inter prediction, or the subdivision into blocks 80 is performed in a different way to the way illustrated in Figure 3. Similarly, a decoder according to an embodiment of the present application may perform block-based decoding of picture 12′ from data stream 14 using the coding concepts further outlined below, but may differ from, for example, decoder 20 of FIG. 2 in that it is a still picture decoder rather than a video decoder, in that it does not support intra prediction or in that it subdivides picture 12′ into blocks in a different way than described with respect to FIG. 3, and / or in that it does not derive prediction residuals from data stream 14 in the transform domain, but rather, for example, in the spatial domain.
[0017] Here, each encoder 60 1 , 60 2 and each decoder 65 1 , 65 2 An embodiment of the present invention will now be described at least in part with reference to Figures 4a and 4b, which respectively illustrate the functionality of the selected inter-component transform 62 of the present invention. 1 or 62 2 , and its reverse version 62 1 ' or 62 2 ' are offset from each other to take into account the order in which they are applied. [Background technology]
[0018] 1. Introduction, State of the Art In natural still and moving color pictures (hereafter simply referred to as images and videos), a significant amount of signal correlation between individual color components can typically be observed. This is especially true for content represented in the YUV or YCbCr (luma-chroma) or RGB (red-green-blue) domains. To efficiently exploit such inter-component redundancy in image or video coding, several prediction techniques have been recently proposed. Among these, the most notable are: Cross-component Linear Model (CCLM) prediction, a Linear Predictive Coding (LPC) method that predicts, at block level, the input signal of one component from the signal of another (usually luma) decoded component and codes only the error, i.e. the difference between the input and the prediction; Joint Chroma Coding (JCC), a technique that encodes only the difference between two chroma residual signals (i.e. only one downmix) and decodes the two chroma signals using a simple sample-by-sample upmix rule "V=-U" or "Cr=-Cb" for YUV or YCbCr coding, respectively. In other words, the JCC upmix represents the prediction of V or Cr from U or Cb, respectively, without coding the associated error or residual for Cr of V, respectively, during the JCC downmix process.
[0019] Both CCLM and JCC techniques, described in detail in [1] and [2], respectively, signal their activation in a particular coded block to the decoder by a single flag. Moreover, both schemes can in principle be applied between any pair of components, i.e. Between a luma and a chroma signal, or between two chroma signals in YUV or YCbCr coding, Between the R and G signals in an RGB coding, or between the R and B signals, or finally between the G and B signals.
[0020] In the above list, the term "signal" may refer to a spatial domain input signal within a particular domain or block of an input image or video, or may represent the residual (i.e., difference or error) between said spatial domain input signal and a spatial domain predicted signal obtained using any spatial, spectral, or temporal predictive coding technique (e.g., angular intra-prediction or motion compensation).
[0021] 2. Shortcomings in the state of technology Although the above solutions have been successful in increasing the coding efficiency in modern image or video codecs, two drawbacks can be identified in relation to the CCLM and JCC approaches: The application of the CCLM method between two chroma channel signals requires, both at the encoder and at the decoder, a computationally relatively complex derivation of certain prediction parameters (CCLM weights) from the upper and left neighboring samples of the considered coded block.
[0022] Using the JCC technique proved to be relatively inflexible, since only signal differences are supported for downmixing and upmixing. On average, the technique works well for YUV or YCbCr coded content, but the coding gain was found to be relatively low for RGB coded inputs, and for natural images or videos recorded with cameras with significant chromatic aberrations. Summary of the Invention [Problem to be solved by the invention]
[0023] It is therefore desirable to provide a more flexible method and apparatus for joint component coding of images or videos that retains the low complexity of the JCC approach. [Means for solving the problem]
[0024] 3. Overview of the Invention To address the above drawbacks, the present invention includes the following aspects, where the term signaling represents the transmission of coded information from an encoder to a decoder. Each of these aspects will be described in detail in a separate section.
Embodiments for Carrying Out the Invention
[0025] 1. The selective application (i.e., activation) of one of the block or picture of at least two inter-component joint coding / decoding methods is carried out together with an on / off flag (optionally entropy-coded) or explicit signaling for each corresponding block or picture of the application of said joint coding / decoding using a non-binary index. Two or more inter-component methods can represent any of the following: · Coding of a single downmix channel representing two color channels, where C’ represents the decoded downmix channel, and the decoded color channels are obtained by Cb’ = a C’ and Cr’ = b C’, where a and b represent specific mixing coefficients (often one of a or b is set equal to 1). · Coding of two mixing channels, where C 1 ’ and C 2 ’ are the decoded mixing channels, and the decoded color components Cb’ and Cr’ are obtained by applying an orthogonal (or nearly orthogonal) transform of size 2 to the decoded mixing channels C 1 ’ and C 2 ’.
[0026] Both methods can be extended to three or more color components. When mixing is applied to N > 2 color components, it is also possible to code M < N (with M > 1) mixing channels and reconstruct the N color components from the M < N decoded mixing channels.
[0027] 2. When joint coding / decoding is applied (i.e., activated), implicit signaling of one of the applied inter-component methods by the existing coded block flag bitstream elements. 3. Direct or indirect signaling per block or picture of the decoding parameters (e.g. upmix matrices, inverse transform types, inverse transform coefficients, rotation angles, or linear prediction coefficients) of all inter-component joint coding / decoding methods applied in said block or picture; 4. At picture or block level, a fast encoder side decision (instead of an exhaustive search) when selecting one of at least two inter-component joint coding / decoding methods to be applied.
[0028] 3.1.Selective application of ICT through explicit application signaling It is proposed to allow arbitrary and selective application of inter-component transforms (ICT) for joint residual sample coding during image or video encoding. As shown in Fig. 1, this ICT design applies a forward joint component transform (downmix) before or after a conventional component-wise residual transform during coding, and applies a corresponding inverse joint component transform (upmix) after or before a conventional component-wise inverse residual transform during decoding. However, unlike the prior art of Section 1 or Section 2, the encoder is given the possibility to select more than one ICT method during coding, i.e. to apply no ICT coding or to apply one ICT method of a set of at least two ICT methods. Combined with the inventive aspects of Section 3.3, this provides more flexibility than the prior art.
[0029] The selection and application (also called activation) of a particular one of at least two ICT methods can be performed globally for each image, video, frame, tile or slice (also slice / tile in more recent MPEG / ITU codecs, hereafter simply called picture). However, in hybrid block-based image or video coding / decoding, it is preferably applied in a block-adaptive manner. The block for which the application of one of the multiple supported ICT methods is selected can represent either a coding tree unit, a coding unit, a prediction unit, a transform unit or any other block within said image, video, frame or slice.
[0030] Whether and which of the multiple ICT methods are applied is signaled in the bitstream using one or more syntax elements at picture, slice, tile or block level (i.e. at the same granularity as the ICT is applied). In one embodiment (described further in section 3.2), the fact that the inventive ICT coding is applied or not is signaled using a (possibly entropy coded) on / off flag for each of said pictures or for each of the blocks for which the ICT coding is applicable. In other words, the activation of the (at least two) inventive ICT methods is explicitly signaled by a single bit per picture or bin and block (where a bin indicates the entropy coded bits, which can consume an average size of less than one bit with the appropriate coding) for the respective blocks. In a preferred version of this embodiment, the application of the ICT methods is signaled by a binary on / off flag. The information of which of the multiple ICT methods is applied is signaled via a combination of additionally transmitted coded block flags (details continue in section 3.2). In another embodiment, the ICT method and the application of the ICT method used are signaled using non-binary syntax elements.
[0031] For both embodiments, a binary or non-binary syntax element indicating the use of an ICT method may only be present (in the syntax) if one or more coded block flags (indicating whether a transform block has non-zero transform coefficients) are equal to 1. In the absence of an ICT related syntax element, the decoder infers that no ICT method is used.
[0032] Furthermore, the high-level syntax may contain syntax elements that indicate the presence of block-level syntax elements as well as their meaning (see Section 3.3). On the one hand, such high-level syntax elements may indicate whether any of the ICT methods are available for the current picture, slice or tile. On the other hand, the high-level syntax may indicate which subset of a larger set of ICT methods is available for the current picture, slice or tile of a picture.
[0033] In the following, specific variants of inter-component conversion are described. These variants are described for two specific color components on the example of chroma components Cb and Cr of typically used YCbCr format image and video signals. Nevertheless, the invention is not limited to this use case. The invention can also be used for any other two color components (e.g., for the red and blue components of an RGB video). Furthermore, the invention can also be applied to the coding of more than two color components (e.g., the three components Y, Cb, and Cr of a YCbCr video, or the three components R, G, and B of an RGB video, etc.).
[0034] ICT Class 1: Transformation-Based Coding In the first ICT variant, two color channels
number
number
number
number
number
number
number
number
number
number
number
[0035] ICT Class 2: Downmix-based coding with reduced number of color channels As mentioned above, the main advantage of the above transform-based ICT variants is that the variance of one of the resulting components is small compared to the variance of the other components (for blocks with a certain amount of correlation). Often this leads to one of the components being quantized to 0 (for the whole block). To simplify the implementation, the color transform is
number
number
number
number
number
number
[0036] Two or more supported ICT methods can include zero or more variations of the transform-based method (specified by a rotation angle or scaling factor) and zero or more variations of the downmix-based method (specified by a rotation angle or scaling factor, and in some cases having an additional flag that specifies which color components are set equal to the transmitted components). This includes cases where (a) all ICT methods represent a transform-based variation, (b) all ICT methods represent a downmix-based variation, and (c) two or more ICT methods represent a mixture of transform-based and downmix-based variations. It should be noted again that the rotation angle or mixing coefficient is not transmitted in block units. Instead, a set of ICT methods is predefined and known to both the encoder and the decoder. In block-based, only an index identifying one of two or more ICT methods is signaled (by a binary flag or non-binary syntax element). A subset of the predefined set of ICT methods may be selected on a sequence, picture, tile, or slice basis, in which case the index coded in block-based signals the method selected from the corresponding subset.
[0037] According to one embodiment, the blocks of samples of the color components are transmitted using a transform coding concept, consisting or at least including a 2d transform mapping the blocks of samples to blocks of transform coefficients, a quantization of the transform coefficients and an entropy coding of the resulting quantization indices (also called transform coefficient levels). At the decoder side, the blocks of samples are reconstructed by first inverse quantizing the entropy decoded transform coefficient levels to obtain reconstructed transform coefficients (inverse quantization typically consists of a multiplication with a quantization step size) and then applying an inverse transform to the transform coefficients to obtain the blocks of reconstructed samples. Furthermore, the blocks of samples transmitted using transform coding often represent a residual signal specifying the difference between the original signal and a prediction signal. In this case, the decoded blocks of the image are obtained by adding the reconstructed blocks of residual samples to the prediction signal. At the decoder side, the ICT method can be applied as follows:
[0038] An ICT transform is applied to the reconstructed transform coefficients (after inverse quantization), then the ICT transform is followed by an inverse 2d transform of the individual color components and, if applicable, the addition of a prediction signal; · An ICT transform is applied to the reconstructed residual signal. This means that the coded color components are first inverse quantized and then inverse transformed by a 2d transform. The resulting block(s) of residual samples are transformed using an ICT transform and a prediction signal may be added after the ICT transform.
[0039] Note that if both ICT and 2d transforms do not include rounding, then both of these configurations will yield the same results. However, in embodiments, all transforms can be specified in integer arithmetic, including rounding, so the two configurations yield different results. Note that it is also possible to apply the ICT transform before inverse quantization and / or after prediction signal addition.
[0040] As mentioned above, the actual implementation of an ICT method may deviate from a unitary transformation (due to the introduction of scaling factors that simplify the actual implementation). This fact should be taken into account by modifying the quantization step size accordingly. That is, in one embodiment of the present invention, the selection of a particular ICT method implies a particular modification of the quantization parameters (and thus the resulting quantization step size). The modification of the quantization parameters may be realized by a delta quantization parameter, which is added to a standard quantization parameter. The delta quantization parameter may be the same for all ICT methods or different delta quantization parameters may be used for different ICT methods. The delta quantization parameters used in connection with one or more ICT methods may be hard-coded or may be signaled as part of a high-level syntax for a slice, a picture, a tile, or a coded video sequence.
[0041] 3.2. Applied implicit signalling of at least one of two ICT methods As mentioned in section 3.1, the activation of one of the at least two ICT methods of the present invention is preferably explicitly signaled from the encoder to the decoder using an on / off flag to instruct the decoder to apply the backward ICT (i.e., the transpose of the ICT processing matrix) when decoding. However, it is still necessary to inform the decoder which of the at least two ICT methods is applied to the processed picture or block at hand for each picture or block for which ICT coding (i.e., forward ICT) and decoding (i.e., backward ICT) are active. Intuitively, an explicit signaling of the particular ICT method (using one or more bits or bins per picture for the respective block) could be used, but it has been found that this form of signaling minimizes the side information overhead of the ICT scheme of the present invention, so implicit signaling is preferably used.
[0042] There are two preferred embodiments of the implicit signaling of the applied ICT method. Both make use of the existing "residual zeroness" indicator in modern codecs such as HEVC and VVC [3], specifically the Coded Block Flags (CBF) bitstream elements associated with each color component of each transform unit. A CBF value of 0 (false) means that the residual block is not coded (i.e. all residual samples are quantized to 0 and therefore there is no need to transmit the quantized residual coefficients in the bitstream), whereas a CBF value of 1 (true) means that at least one residual sample (or transform coefficient) is quantized to a non-zero value for a given block and therefore the quantized residual of said block is coded in the bitstream.
[0043] 3.2.1. Implicit signalling of one of two ICT methods In case of joint ICT coding of a two-component residual signal, two CBF elements are available for implicit ICT method signaling. In case of providing two ICT downmix / upmix methods, the preferred implicit signaling is as follows: [Table 1]
[0044] 3.2.2. Implicit signalling of one of three ICT methods If two CBF elements are available for implicit ICT method signaling as in subsection 3.2.1, but three ICT downmix / upmix methods are provided instead of two for the application, the preferred implicit signaling is as follows: [Table 2]
[0045] If the CBF of both color components in a block is zero, no non-zero residual samples are coded in the bitstream for either component and there is no need to signal information about the applied ICT method.
[0046] 3.3. Optional Direct or Indirect Signaling of ICT Decoding Parameters In the previous section, we described how the activation of an ICT method in a picture or block is explicitly signaled (using an on / off flag) and how the actual selection of one of at least two ICT methods for an affected color component is implicitly signaled (by the existing CBF "residual zeroness" indicator). The set of two or more possible ICT methods can include size-2 Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) or Walsh-Hadamard Transform (WHT) or Karhunen-Loeve Transform (KLT, also known as Principal Component Analysis, PCA) instances, or a predefined (fixed) or input-dependent (adaptive) parameterization of a Givens rotation or linear predictive coding function. All these ICT methods result in one or two downmix signals given two input residual signals in forward form, and two upmix signals given one or two (possibly quantized) downmix signals in backward realization.
[0047] The set of two or more ICT methods with a fixed parameterization can be characterized, for example, by a particular pre-selection of a size 2 transform or a rotation angle or coefficients of a linear predictor function. This parameterization does not need to be transmitted in the bitstream, since it is known to both the encoder and the decoder. In the prior art [2], a fixed "-1" parameterization is used, which results in a downmix rule "C=(Cb-Cr) / 2" and an upmix rule "Cb'=C, Cr'=-C". In the present approach, two or more ICT methods are available for selection by the encoder, and the fixed set of two ICT methods (see section 3.2.1) is [Table 3]
[0048] On the other hand, a fixed set of three ICT methods (see subsection 3.2.2) may be preferable to a set of two; [Table 4]
[0049] This fixed set of 3 ICT design is similar to the sum-difference coding technique commonly applied in both perceptual and lossless audio coding [4,5] and provides significant coding gain. However, this fixed approach was found to result in a relatively non-uniform distribution of said coding gain across the two processed component signals. To compensate for this issue, a more general rotation-based approach can be pursued, realized using a size-2 KLT, also known as Principal Component Analysis (PCA). In this case, the downmix rule is C 1 = Cb cosα + Cr sinα or C 1 = Cb sinα + Cr cosα, C 2 =-Cb sinα+Cr cosα or C 2 = Cb cosα - Cr sinα, This in this case represents a forward KLT over the two components, and the respective upmix rules are Cb'=C 1 '·cosα-C 2 '·sinα or Cb'=C 1 '·sinα+C 2 '·cosα, Cr'=C 1 '·sinα+C 2 '·cosα or Cr'=C 1 '·cosα-C 2 '·sinα,
[0050] It therefore represents the backward KLT, see also [6]. Note that for the rotation angle α=π / 4, the notation on the right hand side of the above equation represents an orthogonal version of the third (cubic) ICT method of the fixed set of three ICT methods in the above equation. In the KLT / PCA approach, the individual first-order, second-order, and optionally third-order ICT methods above can be parameterized using different values of the rotation angle -π≦α≦π. Specifically, α 1 =-π / 8, α 2 =π / 8, possibly α 3 A fixed angle such as α = -π / 4 may be defined for the set of three ICT methods, 1 , α 2 , α 3 is known to both the encoder and the decoder. A single output component variant of the KLT / PCA downmix rule may be defined, C 1 '=0 or C 2 '=0, so the upmix rule is simplified to the coded C 1 ' only or coded C 2 It is worth noting that it is possible to reconstruct the Cb' and Cr' component signals from the 'signal alone (see section 3.1). In this way, a fully flexible and generalized set of two or more ICT methods is constructed that can include the above two and three sets of fixed ICT parameterizations as subsets. This concludes the fixed parameterization aspect.
[0051] It should be noted that for the image and video coding domain, usually only the bitstream syntax and the decoding process are specified. In that context, the described downmix (forward ICT transform) should be interpreted as a specific example for obtaining the downmix channel for a specific upmix rule. The actual implementation in the encoder may deviate from these examples.
[0052] In some coding configurations, it is beneficial to determine the rotation angle α in an input-dependent adaptive manner. In such a scenario, α is determined from two input component signals (here the Cb and Cr residuals) as follows: α=1 / 2 tan -1 (2 CbCr / (Cb 2 -Cr 2 )) or α=1 / 2 tan -1 (2 CbCr / (Cr 2 -Cb 2 )) It can be calculated according to the applied notation of the KLT downmix / upmix rule (see previous page). The above method of deriving α is based on a correlation-based (i.e. least squares) approach. Alternatively, the formula, α = sign(CbCr) tan -1 (sqrt(Cr 2 ) / sqrt(Cb 2 )) or α = sign(CbCr) tan -1 (sqrt(Cb 2 ) / sqrt(Cr 2 ))of, Again, depending on the particular KLT downmix / upmix notation that can be used, this calculation represents the intensity-based principle angle calculation. Both the correlation-based and intensity-based derivation methods (which yield nearly identical results for natural image or video content) utilize the dot product, CbCr=sum b∈B (Cb b ·Cr b ), Cb 2 =sum b∈B (Cb b ·Cb b ), Cr 2 =sum b∈B (Cr b ·Cr b ),
[0053] where B is equal to the set of all sample positions that belong to the processed coded block (or picture). -1is typically implemented using the atan2 programming function to obtain α with sign that is correct, i.e., in the appropriate coordinate quadrant. The derived -π≦α≦π may be quantized (i.e., mapped) to one of a predetermined number of angles and transmitted to the decoder at block or picture level along with the ICT on / off flag. In particular, the following transmission options may be used to inform the decoder about the particular parameterization to apply during backward ICT processing:
[0054] · First option: for each coding block and / or each ICT method used in that coding block, transmit the quantized / mapped α of that ICT method either directly as a quantized angle value or indirectly as an index into a look-up table for the given angle. If only one ICT method is applied to a block and a quantized / mapped α is transmitted for each block, only one α is transmitted. If no ICT coding is active in a block, no quantized / mapped α is transmitted for this block for efficiency.
[0055] · Second option: Send the quantized / mapped α values once per picture or video (set of pictures) for all ICT methods applied or applicable in said picture or video. This can be done at the beginning of the picture or video, for example in the picture parameter set or preferably in the slice header for HEVC or VVC [3]. If no ICT coding is active in the picture or video and / or no chroma coding is performed (e.g. luma only input), there is no need to send the quantized / mapped α values. Again, each α parameter can be sent directly as a quantized angle value or indirectly as an index into a look-up table of predefined angle values.
[0056] Both options can be combined either in parallel or sequentially. To conclude the discussion of adaptive parameterization aspects, it should be noted that it is clear to those skilled in the art that slight deviations from the above parameter transmission options are easily feasible. For example, the ICT parameter transmission per picture or block from the encoder to the decoder may be performed only for selected ICT methods of a set of two or more ICT methods available for coding, e.g., only for methods 1 and 2, or only for method 3. Furthermore, it is clear that in the case of a transform size of 2 (i.e. ICT over two color components), the KLT is equivalent to a DCT or a WHT with α=π / 4 or α=−π / 4. Finally, other transforms than the KLT or, in general terms, downmix / upmix rules may be used as ICT, which may be subject to other parameterizations than the rotation angle (in the most general case, the actual upmix weights can be quantized / mapped and transmitted).
[0057] 3.4. Acceleration Encoder Side Selection of Applied ICT Method In modern image and video encoders, one of several supported coding modes is usually selected based on a Lagrangian bit allocation technique. That is, for each supported mode m (or a subset thereof), the resulting distortion D(m) and the resulting number of bits R(m) are calculated, and the mode that minimizes the Lagrangian function D(m)+λ R(m), where λ is a fixed Lagrangian multiplier, is selected. The encoder complexity grows with the number of supported modes, since the determination of the distortion term D(m) and the rate term R(m) typically requires a 2d forward transform, a (rather complex) quantization, and a test entropy coding for each mode. Therefore, the encoder complexity also grows with the number of ICT modes supported on a block basis. However, there are possibilities to reduce the complexity of the encoder for evaluating ICT methods. In the following, three examples are highlighted:
[0058] · In the encoder, an optimal rotation angle α can be derived based on the original (residual samples) of the color components of the block (e.g., by one of the methods described above). Then, given the derived angle, only the ICT method representing the rotation closest to this angle is tested by deriving the actual distortion D(m) and the actual number of bits R(m) required for this method m.
[0059] · When only a downmixing method is supported (i.e., a method in which N color components are represented by M < N transmission channels), the distortion caused only by downmixing can be evaluated. Next, only the method m that results in the minimum downmixing distortion is tested using the Lagrange technique (i.e., by deriving the actual distortion D(m) and the actual bit rate R(m) associated with method m).
[0060] · When coding two mixed channels C 1 ’ and C 2 ’, both of these channels require non-zero CBFs as in the case of Method 3 in Sec. 3.2.2. After quantization of the first mixed channel (e.g., C 1 ’), it is possible to speed up the encoder by testing whether the quantized version of the first mixed channel exhibits at least one non-zero quantization coefficient. If so (i.e., its CBF is non-zero), the second mixed channel (e.g., C 2 ’) can be quantized, and then the two-channel method can be tested using the Lagrange method. However, if the quantized version of the first mixed channel shows only zero quantization coefficients (i.e., its CBF is 0), quantization of the second mixed channel can be skipped, and for a given quantization parameter, the two-channel method cannot be implicitly signaled and is therefore prohibited, so the Lagrange test for the two-channel method can be aborted.
[0061] 3.5. Context Modeling for ICT Flags and Modes The signaling of ICT usage can be combined with the CBF information. If both CBF flags, i.e. the CBF of each transform block (TB) of each chroma component, are equal to 0, no signaling is required. Alternatively, depending on the configuration of the ICT application, the ICT flag may be sent in the bitstream. The distinction between internal and external context modeling is useful in this context, i.e. internal context modeling selects a context model in the context model set and external context modeling selects a context model set. The configuration for internal context modeling is, for example, the evaluation of the neighboring TBs, using the above and left neighbors and checking their ICT flag values. The mapping from values to context indexes in the context model set can be additive (i.e. c_idx=L+B), exclusive-or (i.e. c_idx=(L<<1)+A) or active (i.e. c_idx=min(1,L+B)). For external context modeling, the CBF condition of the ICT flag can be used. For example, for a configuration using three transforms differentiated by CBF flag combinations, a separate context set is employed for each of the CBF combinations. Alternatively, both the external and internal context modeling can take into account tree depth and block size, such that different context models or different sets of context models are used for different block sizes.
[0062] In a preferred embodiment of the present invention, a single context model is used for the ICT flag, i.e., the context model set size is equal to 1.
[0063] In a further preferred embodiment of the present invention, the intra-context modeling evaluates adjacent transform blocks to derive a context model index, where the context model set size is equal to 3 when using additive evaluation.
[0064] In a preferred embodiment of the invention, the external context modelling uses a different context model set for each CBF flag combination, resulting in three context model sets when the ICT is configured such that each CBF combination results in a different ICT transformation.
[0065] In a further preferred embodiment of the present invention, the external context modeling uses a dedicated context model set if both CBF flags are equal to 1, and uses the same context model set otherwise.
[0066] The description provided herein with reference to features of the encoder also applies, but is not limited to, to a respective decoder adapted to receive a signal or bitstream directly from the encoder, for example using a data connection such as a wireless or wired network, or indirectly using a storage medium such as a portable medium or a server. Conversely, features described in relation to a decoder may be implemented without limitation as corresponding features of an encoder according to an embodiment. This includes, among other features, that features related to a decoder that rely on directly and explicitly evaluating information disclose respective features of an encoder for generating and / or transmitting the respective information. In particular, the encoder may comprise functionality corresponding to the claimed decoder, in particular for testing and evaluating selected encodings.
[0067] Although some aspects are described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, and that a block or apparatus corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
[0068] The encoded image or video signal of the present invention can be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.
[0069] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The embodiments can be implemented using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, having electronically readable control signals stored therein and cooperating (or capable of cooperating) with a programmable computer system such that the respective methods are performed.
[0070] Some embodiments according to the invention comprise a data carrier having electronically readable control signals which cooperate with a programmable computer system to perform one of the methods described herein.
[0071] Generally, embodiments of the present invention may be implemented as a computer program product having a program code operative to perform one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
[0072] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0073] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0074] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium or a computer readable medium) containing, and having recorded thereon, the computer program for performing one of the methods described herein.
[0075] A further embodiment of the inventive method is therefore a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be arranged to be transferred via a data communication connection, for example the Internet.
[0076] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0077] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0078] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0079] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as illustrative and explanatory of the embodiments herein.
[0080] 4. References [1] K. Zhang, J. Chen, L. Zhang, M. Karczewicz, “Enhanced cross-component linear model intra prediction,” JVET-D0110, 2016, http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=2806
[0081] [2] J. Lainema, “CE7-rel.: Joint coding of chrominance residuals,” JVET-M0305, Marrakech, Jan. 2019. http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=5112
[0082] [3] B. Bross, J. Chen, S. Liu, “Versatile Video Coding (Draft 4),” v. 4, JVET-M1001, Marrakech, Feb. 2019. http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=5755
[0083] [4] J. D. Johnston, “Perceptual Transform Coding of Wideband Stereo Signals,” in Proc. IEEE Int. Conf. Acoust. Speech Sig. Process. (ICASSP), Glasgow, vol. 3, pp. 1993-1996, May 1989.
[0084] [5] J. D. Johnston and A. J. S. Ferreira, “Sum-Difference Stereo Transform Coding,” in Proc. IEEE Int. Conf. Acoust. Speech Sig. Process. (ICASSP), San Francisco, vol. 2, pp. 569-572, Mar. 1992.
[0085] [6] R. G. van der Waal and R. N. J. Veldhuis, “Subband Coding of Stereophonic Digital Audio Signals,” in Proc. IEEE Int. Conf. Acoust. Speech Sig. Process. (ICASSP), Toronto, pp. 3601-3604, Apr. 1991. https: / / www.computer.org / csdl / proceedings / icassp / 1991 / 0003 / 00 / 00151053.pdf
Claims
1. 1. A block-based video decoding method, the method comprising: decoding, from the data stream, a flag indicating that the residual samples corresponding to the first chroma block and the residual samples corresponding to the second chroma block are jointly coded as residual samples corresponding to a jointly coded block; decoding, from the data stream, one or more coded block flags indicating whether at least one residual sample corresponding to the first chroma block is non-zero and whether at least one residual sample corresponding to a second chroma block is non-zero; determining coefficients based on the coded block flags; obtaining the residual sample corresponding to the first chroma block and the residual sample corresponding to the second chroma block based on the residual sample corresponding to the jointly coded block; The method of claim 1, wherein the coefficients are applied to the residual samples of the jointly coded block to obtain the residual samples corresponding to the first chroma block or the residual samples corresponding to the second chroma block.
2. The method of claim 1 , wherein the first chroma block comprises a blue chroma component of a picture block and the second chroma block comprises a red chroma component of the picture block.
3. The method of claim 1 , wherein decoding the flags comprises applying a context model based on the one or more coded block flags.
4. 2. The method of claim 1 , wherein the residual samples corresponding to the first chroma block have the same values as the residual samples corresponding to the jointly coded block, and the residual samples corresponding to the second chroma block have values obtained by multiplying the residual samples of the jointly coded block by the coefficients.
5. 1. A block-based video coding method, the method comprising: encoding a flag into the data stream indicating that the residual samples corresponding to the first chroma block and the residual samples corresponding to the second chroma block are jointly coded as residual samples corresponding to a jointly coded block; encoding, into the data stream, one or more coded block flags indicating whether at least one residual sample corresponding to the first chroma block is non-zero and whether at least one residual sample corresponding to the second chroma block is non-zero, the one or more coded block flags indicating coefficients; and encoding, into the data stream, the residual samples corresponding to the jointly coded block, wherein the residual samples corresponding to one of the first chroma block or a second chroma block include the coefficients applied to the residual samples corresponding to the jointly coded block.
6. The method of claim 5 , wherein the first chroma block comprises a blue chroma component of a picture block, and the second chroma block comprises a red chroma component of the picture block.
7. The method of claim 5 , wherein encoding the flags comprises applying a context model based on the one or more coded block flags.
8. 6. The method of claim 5, wherein the residual sample corresponding to the first chroma block has the same value as the residual sample corresponding to the jointly coded block, and the residual sample corresponding to the second chroma block has a value obtained by multiplying the residual sample of the jointly coded block by the coefficient.
9. An apparatus for block-based video encoding or decoding of jointly coded residual samples, comprising one or more processors configured to perform the method of any one of claims 1 to 8.
10. A non-transitory computer readable digital storage medium having recorded thereon a computer program having a program code which, when executed, performs the method according to any one of claims 1 to 8.