Methods and devices for encoding and decoding signals
Adaptive predictive coding in the transform domain addresses inefficiencies in existing signal coding methods by using linear prediction to determine prediction residuals, achieving efficient coding and decoding for biomedical signals with varying characteristics.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing signal coding methods are inefficient in terms of computational load and coding compression, particularly for biomedical signals with varying harmonic, periodic, and transient characteristics across different channels or frames.
A method for encoding and decoding signals using adaptive predictive coding in the transform domain, which applies linear prediction based on reconstructed samples to determine prediction residuals, allowing efficient coding and decoding without requiring side data, suitable for low bitrates and applicable to single- and multi-channel signals.
The method achieves efficient coding and decoding with reduced computational load and improved compression, particularly for signals with transient characteristics, while maintaining low bitrates and enabling frame-by-frame processing.
Smart Images

Figure EP2025081183_07052026_PF_FP_ABST
Abstract
Description
METHODS AND DEVICES FOR ENCODING AND DECODING SIGNALSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 713,216, filed on 29 October 2024 and European Application No. 24209441.5 filed on 29 October 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to encoding and decoding, more specifically to methods and devices for encoding an input signal and methods for decoding an encoded bit stream.BACKGROUND
[0003] Signal coding is frequently used to facilitate efficient transfer, storage and / or processing of single- or multi-channel data signals of various types, such as biomedical signals and audio signals. These, as well as other types of waveform domain signals, may have various characteristics, such as harmonic and periodic characteristics, or transient characteristics (e.g., single pulse or multiple pulses).GENERAL DISCLOSURE
[0004] In some coding applications, such as coding of biomedical signals, the characteristics of the signal waveforms may be different for different types of biomedical signals. The signal characteristics may also vary between frames of a channel or between different channels of a multi-channel signal. For such applications, there is hence a need for coding approaches that are efficient in terms of computational load and coding compression for harmonic and periodic, as well as transient signal waveforms.
[0005] According to a first aspect, there is provided a method for encoding an input signal. The method comprises obtaining a transform domain representation of a channel of a frame of the input signal. The transform domain representation includes a set of input transform domain samples. Each input transform domain sample represents a respective band of the transform domain representation. The method further comprises applying a predictive coding operation to the channel of the frame. The predictive coding operation comprises sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample. The predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands. Theprediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample. The reconstructed sample is determined based on the predicted sample and the prediction residual. The method further comprises forming an encoded bitstream comprising encoded representations of each prediction residual.
[0006] The method of the first aspect provides an adaptive predictive coding approach, applicable to both single- and multi-channel signals, that enables efficient coding in terms of computational load and coding compression. By performing the predictive coding operation in the transform domain, coding gains may be achieved for harmonic and periodic input signal waveforms. Also input signal waveforms with transient characteristics (e.g., single or multiple pulses) within a frame may benefit from prediction in the transform domain. By predicting each sample using linear prediction based on a respective set of reconstructed samples, wherein each reconstructed sample is determined based on a predicted sample and its corresponding prediction residual, the predictive coding approach may operate in a closed-loop fashion, taking into account, for each prediction, both predicted samples and prediction residuals determined for previously processed (i.e., coded) bands. A further benefit of the method is that the coding, and the corresponding decoding, does not require side data. The method is hence well-suited for low bitrates. Moreover, the method allows a frame to be coded without coding prior frames. The method may hence be applied on a frame-by-frame basis in an efficient manner.
[0007] According to a second aspect, there is provided a corresponding method for decoding an encoded bit stream. The method comprises receiving the encoded bit stream. The encoded bit stream comprises encoded representations of a set of prediction residuals of a transform domain representation of a channel of a frame (e.g., of an input signal encoded in the bit stream). Each prediction residual represents a respective band of the transform domain representation. The method comprises decoding the encoded bit stream to obtain the set of prediction residuals. The method further comprises applying a predictive coding operation to the channel of the frame. The predictive coding operation comprises sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample. The predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands. The reconstructed sample is determined based on the predicted sample and the prediction residual.
[0008] Thereby, an encoded bit stream (e.g., encoded according to the method of the first aspect) may be decoded by an adaptive predictive coding approach, providing the same or corresponding benefits as the encoding method of the first aspect, albeit during decoding.
[0009] In some embodiments, the linear prediction of the predicted sample of each iteration may comprise forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights, wherein an updated set of prediction weights is determined for each band. Hence, the set of prediction weights of the linear prediction may be updated on a per-sample basis, thereby further contributing to the adaptive nature of the predictive coding.
[0010] In some embodiments, the sequential processing of the bands may proceed in order from high to low frequency bands. This may contribute to an improved prediction, and hence smaller prediction residual, for low frequency bands.
[0011] In some embodiments, the linear prediction of the predicted sample for each band may comprise forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights. Further, in some embodiments, an updated set of prediction weights may be determined for each band. By updating the set of prediction weights for each band (e.g., during the sequential processing of the bands of the channel), an error between the input transform domain samples and the reconstructed samples may be reduced and an improved coding gain may be achieved.
[0012] In some embodiments, updating the set of prediction weights for a respective band may comprise adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, by a respective step size. The respective step sizes for each respective band may be determined based on the prediction residual of the previously processed band, and one or more of the reconstructed samples of the respective set of reconstructed samples. In some embodiments, the respective step sizes may be determined using a least mean squares method. Thus, the respective step sizes may be updated so as to minimize or reduce a mean squared error between the input transform domain samples and the reconstructed samples.
[0013] In some embodiments, each respective set of reconstructed samples may comprise a set of reconstructed samples determined for previously processed bands for the channel. Hence, samples may be predicted using an in-channel linear prediction. While applicable to signals with various characteristics, an in-channel linear prediction may be especially useful with signals that have a transient character, such as single or multiple pulses within a frame, such as electrocardiogram (ECG) signals.
[0014] In some embodiments of an encoding method of the first aspect, the channel may be a channel of a set of channels of the frame of the input signal, and wherein the method may further comprise, for each of the further channels of the set of channels, obtaining a respective transform domain representation of the further channel of the frame, the transform domainrepresentation including a respective set of input transform domain samples for the further channel, wherein each input transform domain sample represents a respective band of the transform domain representation. A coding operation, in particular a predictive coding operation may further be applied to the further channel of the frame. The predictive coding operation may comprise sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample. The predicted sample may be determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands. The prediction residual may be determined based on the input transform domain sample and the predicted sample and represent a prediction error between the input transform domain sample and the predicted sample. The reconstructed sample may be determined based on the predicted sample and the prediction residual. The encoded bitstream may thus be formed to further comprise encoded representations of each prediction residual for each further channel.
[0015] Thus, the encoding method may be applied to a multi-channel input signal, wherein for at least some of the channels, a predictive coding operation may be applied.
[0016] Correspondingly, in some embodiments of a decoding method of the second aspect, the channel may be a channel of a set of channels of the frame (e.g., of an input signal encoded in the encoded bit stream), wherein the encoded bit stream comprises encoded representations of a respective set of prediction residuals of a transform domain representation of each of the channels of the frame, wherein each prediction residual of the respective set of prediction residuals represents a respective band of the transform domain representation, and wherein the method may further comprise, for each of the further channels of the set of channels: decoding the encoded bit stream to obtain the respective set of prediction residuals for the further channel. The method may further comprise applying a predictive coding operation to the further channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample. The predicted sample may be determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands. The reconstructed sample may be determined based on the predicted sample and the prediction residual.
[0017] Thus, the decoding method may be applied to a multi-channel signal, wherein coded samples in the form of prediction residuals for each of the channels may be used to determine reconstructed samples for each channel of the originally encoded input signal.
[0018] In the encoding and decoding methods alike, the channels may be subjected to the coding operation in a sequence (e.g., from a first channel to a last channel of the frame of the input signal / encoded signal).
[0019] The coding operation applied to each channel may comprise in-channel linear prediction and / or cross-channel linear prediction.
[0020] In some embodiments, applying an in-channel linear prediction to a respective channel may comprise basing the linear prediction of each respective predicted sample of the respective channel on a respective first set of reconstructed samples determined for previously processed bands for the respective channel.
[0021] In some embodiments, applying a cross-channel linear prediction to a respective channel may comprise basing the linear prediction of each respective predicted sample of the given channel on a respective second set of reconstructed samples determined for previously processed bands of a respective set of the other channels of the set of channels, for a same band as the respective predicted sample.
[0022] According to a third aspect, there is provided a method for encoding an input signal comprising a set of channels, the method comprising, for each of a first frame, f=fi, and a second frame, = 2, of the input signal:obtaining a transform domain representation of the frame / , the transform domain representation including, for each channel, k, of the set of channels, a set of transform domain samples (coefficients), each transform domain sample A / ( / c, n) associated with a respective band, n, of the transform domain representation;applying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples, a set of prediction residuals, and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, n / o, to a final band, n i, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ;a prediction residual, Df(k, n), is determined based on the transform domain sample A (k, ri) and the predicted sample Pf(k, ri) determined for the given band, n, and represents a prediction error between the transform domain sample Ay( / c, ri) and the predicted sample Py (k, ri); anda reconstructed sample, X'f (k, n), is determined based on the predicted sample Pf(k, ri) and the prediction residual Df(k, ri) determined for the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each prediction weight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band nj2,o in the second frame f.the set of prediction weights is initialized with the set of prediction weights used for processing the final band n / i for the given channel k, in the first frame;, wherein the final processed band nu in the first frame / ; is the same band as the initial processed band / 7 / 2. in the second frame / 2; oreach given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame / ;; and the method further comprising encoding representations of the prediction errors (e.g., quantized representations of the prediction errors) determined for each band n and channel k into a bitstream.
[0023] According to a fourth aspect, there is provided a method for decoding an encoded bitstream, the method comprising:decoding the encoded bitstream to obtain, for each of a first frame, / = / ;, and a second frame, / = / 2, a set of prediction residuals for each channel k of a set of channels, each prediction residual, Df(k, n), associated with a respective band, n, of a transform domain representation;and for each of the first frame / ; and the second frame f2‘.applying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, nf,0, to a final band, nf,1, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ; anda reconstructed sample, X' (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) associated with the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each predictionweight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band nj2,o in the second frame f.the set of prediction weights is initialized with the set of prediction weights used for processing the final bandfor the given channel k, in the first frame;, wherein the final processed band ni in the first frame / ; is the same band as the initial processed band nyz.oin the second frame / 2; oreach given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f1.
[0024] According to a fifth aspect there is provided a device (e.g., an encoder or a decoder) comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to the first, second, third or fourth aspect, or any embodiments thereof.
[0025] According to a sixth aspect there is provided a computer program product comprising computer program code portions configured to perform the method according to the first, second, third or fourth aspect, or any embodiments thereof, when executed on a computer processor.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Embodiments of the present disclosure will be described in more detail with reference to the appended drawings.
[0027] Figure 1 is a block diagram of an encoding device.
[0028] Figure 2 is a block diagram of a decoding device corresponding to the encoding device of Fig. 1.
[0029] Figure 3 is a block diagram of a further encoding device.
[0030] Figure 4 is a block diagram of a further decoding device corresponding to the encoding device of Fig. 3.
[0031] Figure 5 is a block diagram of yet a further encoding device.
[0032] Figure 6 is a block diagram of a decoding device corresponding to the encoding device of Fig. 5.
[0033] Fig. 7a-b show, respectively, a high-level block diagram of an encoder system and a decoder system
[0034] Fig. 8 is a flow chart of a method for encoding an input signal.
[0035] Fig. 9 is a flow chart of a method for decoding an encoded bit stream encoded according to the method of Fig. 8.DETAILED DESCRIPTION
[0036] Fig. 1 is a block diagram of an encoding device or system 100 (hereinafter termed encoder 100) implementing a method for encoding an input signal, which in the following will be described with further reference to the flow chart of Fig. 8.
[0037] The encoder 100 obtains at step SI 01 a set of input transform domain samples X (k, ri) of a transform domain representation of a channel k of a frame of the input signal to be encoded. Each input transform domain sample represents a respective band or bin of the transform domain representation, where n is the bin index of the respective band. The terms band and bin may for convenience in the following be used interchangeably, and each be referenced by the index n. The operations of the encoder 100 will be described mainly with reference to a single channel k, which means that the channel index for the purpose of the following discussion will be assumed to be fixed (e.g., k = 0), or simply be ignored.
[0038] The input signal may be a waveform domain signal comprising a sequence of waveform domain samples, wherein the set of input transform domain samples X(k, ri) obtained by the encoder 100 may be samples X(k, ri) of the transform domain representation of (a channel of) a frame of the waveform domain input signal. For example, the input signal may be a time domain signal comprising a sequence of time domain samples, wherein the set of input transform domain samples X(k, ri) obtained by the encoder 100 may be samples X(k, ri) of the transform domain representation of (a channel of) a frame of the time domain input signal. The time domain input signal may be a time domain biomedical signal comprising one or more channels of time domain biomedical signal samples. A biomedical signal in the context of the present disclosure may be a signal that relates to physiological information, and the signal may be electrical, physical or biochemical. The biomedical signal may relate to biological systems and conditions, examples of which may include Electrocardiography (ECG) data, Electroencephalography (EEG) data, Electromyography (EMG) data, and Photoplethysmogram (PPG) data, or signals for blood sugar level, heart rate, body temperature, respiratory rate and oxygen saturation. Further, a biomedical signal may comprise one or more channels of time domain biomedical signal samples. In other examples, the biomedical signals may relate to muscle and / or skin measurements. Any other medical signal and / or physical response would also be understood to be comprised by this definition. As another example, a time domain input signal may be a time domain audio signal comprising a sequence of time domain audio samples. However, the present disclosure is not limited to time domain input signals, but may moregenerally be applied to encode also other types of input signals that may be transformed into a transform domain representation. More generally, the input signal may be a waveform domain input signal and comprise a sequence of waveform domain samples distributed in a spatial dimension, with each sample indicating a property at a position in space. For example, an input signal may represent one or more channels of an image, with each sample representing a value (e.g. color, brightness, etc.) associated with an individual pixel of the channel of the image.
[0039] Regardless of the initial domain of the input signal (be it a time domain input signal or another waveform domain input signal), a transform T may be applied to the sequence of samples of the input signal (e.g., for the single channel, or for each channel k), to obtain the set of samples X(k, ri) of the transform domain representation (e.g., for the single channel, or a respective set of samples X k, ri) for each channel k).
[0040] The transform T may be applied to the input signal by a transform unit of the encoder 100 (not shown in Fig. 1). Examples of suitable transforms T include frequency transforms such as discrete Fourier or trigonometric transforms, for instance a discrete cosine transform (DCT), such as DCT Type II or Modified DCT (MDCT), or a discrete sine transform (DST), such as DST Type II. For example, the transform T may be an integer discrete transform, such as an integer DCT Type II or an integer DST Type II. As yet another example, the transform T may be realized using an invertible integer discrete Fourier transform as disclosed in U.S. Provisional Patent Application No. 63 / 713,197 titled “Methods and Systems for Integer Invertible Discrete Trigonometric Transforms”, filed by Applicants Dolby Laboratories Licensing Corporation and Dolby International AB on October 29, 2024, which is herewith incorporated by reference in its entirety.
[0041] The number of transform domain samples of the set of transform domain samples X(k,ri) (i.e., the number of bands / bins) is herein denoted N, where the highest band / bin index is n = N — 1 and the lowest band / bin index is n = 0. The number N is given by the transform length of the transform T. The specific choice of transform length may vary depending on the specific application and type of input signal, but may for instance be 256, 512, 1024, 2048, 4096, or greater, to mention a few non-limiting examples. A frame of the input signal may comprise one or more transform blocks, depending on the number of samples of the frame and the transform length. The set of transform domain samples referred to in the following may thus refer to a set of transform domain samples of a transform block of the frame. Hence, where a frame comprises more than one transform block, the encoding and decoding methods described in the following may comprise sequentially processing the transform blocks of the frame. For instance, transform domain samples of a transform domain representation of a channel may referto transform domain samples of a transform block of the channel. The encoding or decoding of a frame may be complete when each transform block of the frame has been encoded or decoded.
[0042] At step S102, the encoder 100 applies a predictive coding operation to the channel k (e.g., the single channel or each channel k). The predictive coding operation comprises sequentially processing each bin of a set of the bins n (i.e., bands n) of the transform domain representation X(k, ri) (e.g., the set of bins n of a transform block of the frame), to determine, for each of the bins of the set, a predicted sample P(k, ri), a prediction residual D(k, ri), and a reconstructed sample X'(k, ri).
[0043] The prediction residual D (k, ri) is determined based on the transform domain sample X(k, ri) and the predicted sample P(k, ri) and represents a prediction error between the input transform domain sample X(k, ri) and the predicted sample P(k, ri). The reconstructed sample X'(k, ri) is determined based on the predicted sample P(k, ri) and the prediction residual D(k, ri). As further described below, the prediction residuals D(k, ri) represent coded samples that will be encoded into the bit stream B output of the encoder 100.
[0044] The predicted sample P(k, ri) is determined by linear prediction (LP) block 110 using linear prediction based on a respective set of reconstructed samples determined for previously processed bins n. In the illustrated example, the LP block 110 implements an in-channel linear prediction, meaning that the LP block 110 determines the predicted sample P(k, ri) using linear prediction based on a respective set of reconstructed samples X’(k, n) determined for previously processed bins n of the same (e.g., single) channel k. The number of reconstructed samples of each respective set of reconstructed samples corresponds to the prediction order U (e.g., in- frame or in-transform block) of the linear prediction. The prediction order U may for example be set to a fraction of the transform length N, such as U = N / R (e.g., rounded to nearest integer value), where 10 < 7? < 100. As a specific non- limiting example the prediction order U may be 40 for a 2048 transform length. For example, with reference to a given transform block of the frame, the respective sets of reconstructed samples may be respective sets of reconstructed samples determined for previously processed bins n of the given transform block of the frame.
[0045] As stated above, the predictive coding operation comprises sequentially processing each bin of a set of the bins n (i.e., bands n) of the transform domain representation X(k, n), to determine, for each of the bins of the set, a predicted sample P(k, n), a prediction residual D(k, ri), and a reconstructed sample X’(k, ri). “The set of the bins n“ may here typically refer to a subset of (typically consecutive) bins of all N bins / bands of the transform domain representation (e.g., of the transform block, as given by the transform length A). This since (assuming only in-channel linear prediction is performed) for the highest frequency bin n = N —1 (in case of “downward linear prediction”, discussed below), or the lowest frequency bin n = 0 (in case of “upward linear prediction”, discussed below), there exists no previously processed bins for which one or more reconstructed samples have been determined. Hence, in case of downward linear prediction, the highest frequency bin n = N — 1 will typically be initialized prior to predicting samples for the lower frequency bins n < N — 1. Correspondingly, in case of upward linear prediction, the lowest frequency bin n = 0 will typically be initialized prior to predicting samples for the higher frequency bins n > 0.
[0046] Therefore, when stated herein that the predictive coding operation comprises sequentially processing the bands / bins of the transform domain representation to determine “for each of the bands / bins”, a predicted sample, a prediction residual and a reconstructed sample, the reference to “each of the bands / bins” refers to the set (i.e., subset) of all N bins / bands of the transform domain representation (e.g., of the transform block, as given by the transform length N ) for which a predicted sample, a prediction residual and a reconstructed sample is determined. This set (subset) may, as explained above, be defined by each of the bands / bins of the transform domain representation except for the highest frequency band / bin (i.e., n = 0... N — 2, in case of downward prediction) or the lowest frequency band / bin (i.e., n = 1... N — 1, in case of upward prediction). In this case, the predictive coding operation may further comprise initializing a highest or lowest frequency band / bin of the N bands / bins of the transform domain representation of the input signal. In other words, with reference to a set including also the highest or lowest frequency band / bin of the N bands / bins of the transform domain representation, the predictive coding operation may comprise sequentially processing each band n of the N bands to:determine or initialize, for a first processed band being the highest or lowest frequency band of the N bands, a “first reconstructed sample” (interchangeably “initialized reconstructed sample”) to correspond to the input transform domain representation of the first processed band, andfor each of the further bands of the set of N bands, determine a predicted sample, a prediction residual, and a reconstructed sample, wherein each respective predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for bands previously processed in the predictive coding operation. Thus, for the first U bands of the further bands, the linear prediction may also be based on the first / initialized reconstructed sample. Thus, the “first reconstructed sample” X'(k, n0) may be set to be equal to the input transform domain representation X(k, n0), or be equal to Q-1{Q{X(k, n0)}}, where n0denotes the highest or lowest frequency bin, e.g., for downward prediction, n0= n — 1 and for upward prediction n0= 0. Q {} and Q1{} represent a quantization operation and an inverse quantization operation, and are further discussed below.
[0047] As a further example, it is possible to extend this initialization approach beyond the (single) highest or lowest frequency band of the N bands. Thus, according to a more general example, the predictive coding operation may first process each band of the set of L > 1 highest or lowest frequency bands of the N bands to determine or initialize, for the set of L bands, a set of L reconstructed samples (i.e., a “first” I “initialized” set of L reconstructed samples) to correspond to the respective input transform domain representations of the set of L bands. Thus, for each band n0of the set of L bands the reconstructed sample X'(k, n0) may be set to be equal to the input transform domain representation X(k, n0), or be equal to Q-1{Q{X(k, n0)}}.Subsequently, the further or remaining N — L bands may be sequentially processed as set out above, i.e. to, for each of the further bands of the set of N bands, determine a predicted sample, a prediction residual, and a reconstructed sample, wherein each respective predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for bands previously processed in the predictive coding operation, which thus may include one or more reconstructed samples of the set of L bands (being previously processed / initialized bands).
[0048] The linear prediction of a predicted sample P(k, n), as implemented by the LP block 110, may comprise forming a weighted combination of its respective set of reconstructed samples using a set of prediction weights. For example, the LP block 110 may determine a predicted sample P(k, n) = PLP(k, ri) for a channel k and a band n according to:(Eq. 1) where a(n, u) are prediction weights, X' (k, n + 1 + u) is a reconstructed sample of the respective set of reconstructed samples, and U is the number of reconstructed samples (i.e., the in-frame prediction order). As the prediction is performed in the transform domain and is based on reconstructed samples of previously processed bands, the prediction may be referred to as a cross-band prediction.
[0049] As may be seen from Eq. 1, the sequential processing of the predictive coding operation (i.e., the prediction) may proceed in order from high to low frequency bands / bins n. That is, the prediction may start at n = N — 1 (the highest frequency bin index) and proceed until n = 0 (the lowest frequency bin index). The predictive coding operation according to Eq. 1 may thus be referred to as a downward linear prediction.
[0050] A reconstructed sample X'(k, n) determined for a bin n corresponds to a reconstruction of the transform domain input sample X (k, ri) of the same band (same bin indexn). As shown in Fig. 1, the encoder 100 may determine the reconstructed sample X'(k, n) according to:X'(k,n) = D(k,ri) + P(k,ri)(Eq. 2) by the sum block 108 summing the predicted sample P(k, n) output by the LP block 110 and the prediction residual D (k, ri).
[0051] As further shown in Fig. 1, the prediction residual D(k, ri) may be determined by determining, at difference block 102, a difference between the transform domain input sample X(k, n) and the predicted sample P(k, n), and subsequently quantizing (at quantization block Q 104) and then inverse quantizing (at inverse quantization block iQ 106) the quantized difference. Thus, the prediction residual D(k, ri) is given by the output from the iQ block 106 according to:(Eq. 3) where Q {} corresponds to the quantization operation performed by the Q block 104, and Q-1{} corresponds to the inverse of the quantization operation of the Q block 104 performed by the dequantization block 106. The Q block 104 may implement any suitable quantization operation Q {}. The iQ block 106 may implement the corresponding inverse quantization operation Q-1{}, such as multiplying Q{X(k, n) — P(k, n)} with the quantization factor of Q{}. For sake of completeness, it is noted that in case a quantization block 104 is omitted, or the quantization block 104 is arranged after the branching point B indicated in Fig. 1, the prediction residual D (k, ri) may simply be given by:D(k, ri) = X(k, ri) — P (k, ri)(Eq. 3’)
[0052] Each reconstructed sample X'(k, n) and prediction residual D(k, n) is stored (e.g., buffered) in sample buffer 114. Optionally, the sample buffer 114 may also store the transform domain input samples X (k, ri).
[0053] Prior to initiating the linear prediction, the reconstructed samples X'(k, m) may be initialized asX'(k, m) = 0 for m > N — 1(Eq. 4) Again, in the case of a single-channel input signal, the channel index k may be ignored. In any case, this initialization allows the LP block 110 to use a fixed prediction order U throughout the sequence, although U number of reconstructed samples may not be available until after processing U bins.
[0054] Further, as may be seen from Eq. 1, following this initialization, for the first processed band n = N — 1, the predicted sample P(k, n = N — 1) will become 0, wherein the prediction residual D(k, n = N — 1) becomes D(k, N — 1) = Q-1{Q{X(k, N — 1)}} (according to Eq. 3) or D(k, N — 1) = X(k, N — 1) (according to Eq. 3’). Hence, with this initialization, Eq.1 is defined for each of the bins n = 0... N — 1 and may thus be applied for processing each of the bins n = 0... N — 1. However, in light of the preceding discussion concerning the initialization, it is noted that the first bin n for which a predicted sample P(k, ri) is determined based on a reconstructed sample X'(k, m) determined for a previously processed bin m = n + 1 as D(k, m) + P(k, m), is n = N — 2.
[0055] As an alternative to this initialization, the summation range for u in Eq. 1, may be shifted during the course of the sequential processing to include only existing frequency bands / bins within the transform length N.
[0056] The set of prediction weights used by the LP block 110 to form a weighted combination of reconstructed samples (e.g., a(n, u) of Eq. 1), may be determined by Coefficient update block 112. For example, denoting like in Eq. 1 a prediction weight for a band n by a(n, u), the set of prediction weights for band n is given by a(n, 0)... a(n, U — 1). The Coefficient update block 112 may determine an updated set of prediction weights for each band n. For example, the set of prediction weights a(n — 1,0)... a(n — 1, U — 1) for a band n — 1 may be updated by adjusting each prediction weight of the set of prediction weights a(n, 0)... a(n, U — 1) used for its previously processed neighboring band n by a respective step size. That is, an updated prediction weight for a given band (e.g., a(n — 1, u) for band n — 1) may be determined based on its corresponding prediction weight used for its previously processed neighboring band (e.g., a(n, u) for band n) and a respective step size. The set of prediction weights may be stored in an internal buffer of the Coefficient update block 112, or in the sample buffer 114.
[0057] A goal of the prediction may be to minimize the energy of the prediction residuals D (k, ri). This may be achieved by updating the set of prediction weights using a least mean squares (LMS) method, for instance a normalized LMS (NLMS) method. More specifically, the respective step size for adjusting each respective prediction weight may be determined using an LMS or NLMS method. For example, the Coefficient update block 112 may determine updated prediction weights according to:a(n — 1, u) = a(n, u) + D(k, ri) g X'(k,n + 1 + u)(Eq. 5) where g is a gain factor given by:(Eq. 6) where r < 1 is a factor which may be set in dependence on the transform length N. For example, r = 0.25 may be a good choice for typical transform lengths N, e.g., N = 2048 as a non- limiting example. ||X'(k, n)|| denotes the norm of the respective set of reconstructed samples ||X'(k, n)|| determined for the N previously processed bins n. The norm may be computed as:(Eq. 7)
[0058] In some implementations, the computation of r / ||X'(k,n)|| in Eq. 6 may efficiently be ||X'(k,n)|| approximated by a right bit-shift by ceil{log2||X'(k, n)||2}+1 bits. That is, r / ||X'(k,n)|| in Eq. 6 may be approximated by >> ceil{log2||X'(k, n)||2} >> 1.
[0059] To more fully appreciate the benefits of updating the prediction weight according to Eq. 5 and 6, consider the squared prediction error E given by:E = D(k,n)2(Eq. 8) To find the adjustment amounts (i.e., the respective step sizes), the squared prediction error E may be minimized for a given change of the adjustment amounts. Denoting the adjustment amounts for adjusting the prediction weights a(n — 1, u) for a given band n — 1 by h(n — 1, u), minimizing the squared prediction error E may involve determining a gradient of the squared prediction error E with respect to the adjustment amounts h(n — l,u). The given change of the respective step sizes for the given frequency band may be a change in a sum or squared sum of the adjustment amounts, for example. A maximum reduction of the squared prediction error E for a given change in the squared sum of the adjustment amounts h(n — 1, u)2may be obtained in the opposite direction of the gradient of E with respect to h. This may amount to using adjustment amounts given byh(n — 1, u) = gD(k, n)X'(k, n + 1 + u)(Eq. 9) with a positive gain g. As the prediction target is a scalar, a vanishing squared prediction error E = 0 may be obtained by adjusting g to gfull= 1 / ||X'(k, n)||2. Using g = rgfull= r / ||X'(k, n)||2and setting r < 1, and further limiting g such that the sum of squares of theadjustment amounts h(n — 1, u) is less than r2gives Eq. 6, and may achieve a good trade-off with regard to balance between update speed and quality of prediction.
[0060] The predictive coding operation for the channel k of step SI 02 may be concluded after each of the bins N have been processed. The prediction residuals D(k, n) obtained thereby may represent coded samples for all bins N of the channel k, except for the bin n0, for which the above-mentioned “first reconstructed sample” X'(k, n0) may represent the coded sample. Thus, at step S103 the encoder 100 proceeds to form an encoded bitstream B comprising encoded representations of each of the coded samples.
[0061] Forming the encoded bitstream B may as shown in Fig. 1 comprise quantizing the coded samples at Q block 104, and thereafter entropy encoding each of the (quantized) coded samples by Entropy coding block 116. The Entropy coding block 116 may apply Huffman coding or Golomb-Rice coding. As yet another example, the Entropy coding block 116 may apply the coding scheme as disclosed in U. S. Provisional Patent Application No. 63 / 713,209 titled “Entropy Coding Combining Huffman and Golomb-Rice Coding”, filed by Applicant Dolby Laboratories Licensing Corporation on October 29, 2024, which is herewith incorporated by reference in its entirety.
[0062] While the entropy coding may provide lossless compression of the coded samples, and hence further may reduce a bit rate of the bitstream B, it is also possible to directly include the (quantized) coded samples in the bitstream B without applying entropy coding. The bitstream B may be stored and / or transmitted to a receiving device comprising a decoding system.
[0063] Fig. 2 is a block diagram of a decoding device or system 100’ (hereinafter termed decoder 100’) implementing a method suitable for decoding a bitstream B encoded by the encoder 100, which in the following will be described with further reference to the flow chart of Fig- 9.
[0064] At step S201 the decoder 100’ receives the encoded bitstream B. The decoder 100’ may receive the bitstream B from the encoder 100 via a wired or wireless communication link, or by reading the bitstream B from a memory storing the bitstream B. As explained above, the bitstream B comprises encoded representations (e.g., quantized and / or entropy encoded representations) of the coded samples generated by the encoder 100 from the transform domain representation of the (single) channel k of the frame of the input signal. Thus, the coded samples may include the (quantized) prediction residuals D(k, ri) for all bins N of the channel k, except for the bin or bins n0, for which the above-mentioned “first reconstructed sample” X'(k, n0), or “first set of L reconstructed samples”, may represent the coded sample(s). As mentioned above, where a frame comprises more than one transform block, respective sets of coded samples maybe obtained for each transform block of the frame. The processing described in the below may be applied to each respective transform block.
[0065] At step S202 the decoder 100’ applies at Entropy decoding block 116’ entropy decoding to the encoded bit stream B to obtain the (quantized) prediction residuals D (k, ri) and the first reconstructed sample(s) X'(k, n0). In case no entropy encoding has been applied by the encoder 100, the decoding at step S202 may simply comprise retrieving the coded samples from the bitstream B.
[0066] At step S203 the decoder 100’ applies a predictive coding operation to each of the coded samples to determine, for each of the bands for which a prediction residual D (k, ri) is included in the bitstream B, a predicted sample P k, ri) and a reconstructed sample X'(k, ri). As may be understood in light of the discussion of the initialization for the highest or lowest frequency band n0, no reconstruction is needed for band n0as the first reconstructed sample X' k, n0) (i.e., the coded sample for band n0) is set to X(k, n0) of the initially encoded input signal, or to Q-1{Q{X(k, n0)}}.
[0067] Where the coded samples are quantized, the decoder 100’ may as shown include a corresponding inverse quantization iQ block 106, preceding the stage implementing predictive coding operation.
[0068] The predictive coding operation of the decoder 100’ implements, like the encoder 100, an in-channel predictive coding operation. Hence, the predictive coding operation of the decoder 100’ proceeds analogous to the predictive coding operation of the encoder 100, with the difference that the prediction residuals D (k, ri) here are retrieved from the bitstream B, rather than being derived from predicted samples and input transform domain samples as in the encoder 100. Thus, the LP block 110, the Coefficient update block 112 and the sample buffer 114 may each operate in the same manner as the correspondingly numbered elements of the encoder 100 in Fig. 1. This further means that the predicted samples P(k, n) and the reconstructed samples X'(k, ri) also may be determined according to Eq. 1 and Eq. 2. Further, the reconstructed samples X'(k, m) may be initialized according to Eq. 4. Also, the Coefficient update block 112 may determine updated prediction weights according to Eq. 5 and Eq. 6. The description of the predictive coding operation will therefore not be repeated here, but reference is made to the corresponding passages of the preceding description of the encoding method and the encoder 100. An effect of implementing the decoder 100’ analogous to the encoder 100 is that also the decoder 100’ may derive the prediction weights from the prediction residuals D(k, ri) obtained from the bitstream B. Hence, no prediction weights need to be included in the bitstream B to enable the decoding.
[0069] Hence, for the first bin or bins n0processed in the linear prediction operation, a “first reconstructed sample”, or a “first set of L reconstructed samples” may be determined or initialized to X'(k, n0), as retrieved from the bitstream B. For each further bin n, a predicted sample P(k, n) may be determined (e.g., according to Eq. 1) and a reconstructed sample X'(k, n) may be determined based on the predicted sample P(k, n) and the prediction residual D(k, n). The linear prediction operation proceeds until each of the bins N have been processed, wherein a set of N reconstructed samples X'(k, n) representing the initially encoded channel k of the frame of the input signal have been determined and stored in the sample buffer 114. This set of reconstructed samples X'(k, ri) may be output as a decoded signal from the decoder 100’.
[0070] In the above exemplified implementations of step SI 02 (performed by the encoder 100) and step S202 (performed by decoder 100’), the encoder / decoder 100 / 100’, for each bin n, stores the reconstructed samples X'(k, ri) in the sample buffer 114 and determines the prediction residuals D(k, n) based on the U previously stored reconstructed samples X'(k, l), for l = n + 1... n + U. However, in other implementations of step S102 and / or S202, the sample buffer 114 need not store the reconstructed samples X'(k, ri) during processing of each respective bin n, but may instead store the predicted samples P (k, ri) and the prediction residuals D (k, ri) for each bin n. Thus, the LP block 110 may calculate each reconstructed sample X'(k, n) when needed, e.g., “on-the-fly”. That is, upon determining a given predicted sample P(k, n), the LP block 110 may for each of the U bins, determine its corresponding reconstructed samples X'(k, l) based on the previously stored predicted samples P(k, l) and prediction residuals D(k, l) (e.g., for I = n + 1... n + U).
[0071] In the above, the operations of the encoder 100 and decoder 100’ have been described with reference to a downward linear prediction, proceeding in order from high to low frequency bands / bins n. While it is expected that this approach may enable efficient coding for many types of time domain input signals, such as biomedical and audio signals, it is also possible to sequentially process the bins / bands n in the opposite order, i.e., from low to high frequency bins n. This may be denoted as an upward linear prediction. Hence, for upward linear prediction, counterparts to each of Eq. 1, 4, 5 and 7 may be given by:(Eq. 1’) "(k, m) = 0 for m < 0(Eq. 4’) a(n + 1, ri) = a(n, ri) + D(k, ri) g X'(k,n — 1 — ri)(Eq. 5’)(Eq. 7’)
[0072] Fig. 3 is a block diagram of an encoder 200. The description of the encoder 100 of Fig. 1 is generally applicable to the encoder 200 of Fig. 3, and thus, like reference signs are used in Fig. 1 and Fig. 3 to refer to like elements. For conciseness, a description of such like elements will not be repeated, but rather reference is made to the corresponding discussion of Fig. 1.
[0073] The encoder 200 however differs from the encoder 100 in that instead of an inchannel LP block, the encoder 200 includes a cross-channel linear prediction (LC) block 206.
[0074] To enable the cross-channel linear prediction, the input signal obtained by the encoder 200 is a multi-channel input signal. An encoder 200 using only cross-channel linear prediction may in particular provide coding efficiency when the input signal comprises a relatively large number of channels (e.g., 10 or more). For instance, in some biomedical signal applications, the number of channels may be in excess of 30, and, in some cases, 100 or more. Especially, coding efficiency may be obtained when there are a relatively large number of strongly correlated channels, as the case may be for some electroencephalogram (EEG) signals, as a non-limiting example.
[0075] For the purpose of the following discussion, it will be assumed that a frame of an input signal to be encoded by the encoder 200 comprises a plurality of channels, denoted by M, such as 10 channels or more. In general though, the cross-channel linear prediction as set out in the following may be applied also to input signals including fewer channels. The frame may comprise, for each channel, one or more transform blocks (i.e., depending on the number of samples of the frame and the transform length). Where the frame comprises more than one transform block per channel, the respective transform blocks for the channels are aligned, i.e., such that the respective first transform blocks for each of the M channels are derived from aligned channel portions in the waveform domain (e.g., time-aligned in the time domain), and so on for the subsequent respective transform blocks of the frame.
[0076] Hence, analogous to the encoder 100, at step S 101, the encoder 200 obtains a respective transform domain representation of each channel k of the frame. Each respective transform domain representation includes a respective set of transform domain samples X(k, ri) for the channel k. Each transform domain sample X(k, ri) represents a respective band n of the transform domain representation of the channel k. It is to be noted that the sets of transformdomain samples X(k, ri) of each of the M channels of the frame have the same bin index n, such that transform domain samples with different channel index k but the same bin index n are transform domain samples for a same frequency band. Further, the transform T used to determine the transform domain representation for each channel k has the same transform length N, such that the number of bands / bins for each channel k is N.
[0077] At step SI 02, the encoder 200 applies a coding operation to each of the M channels. For the first channel k = 0, the coding operation is not a predictive coding operation, as there are no previously coded channels to base a predictive coding operation on. Hence, for the first channel k = 0, the coding operation may simply comprise determining or initializing a set of “first reconstructed samples” X'(0,0)... X'(0, N — 1) to correspond to X (0,0)... X(0, N — 1). For example, each first reconstructed sample X'(0, ri) may be set to be equal to the input transform domain representation X(0, n), or be equal to Q-1{Q{X(0, n)}}, where Q{} and Q-1{} like in the above represent the quantization operation of Q block 104 and the inverse quantization operation of iQ block 106, respectively. The thus obtained first reconstructed samples are stored in a channel buffer 202 as a first set of reconstructed samples for a first coded channel k = 0. This type of coding operation applied to the first channel k = 0 may herein be referred to as a non-predictive coding operation.
[0078] For each further channel k > 0 of the frame, a predictive coding operation may be applied. The further channels k > 0 are processed sequentially, i.e., one after another. For each channel k > 0, predictive coding operation is applied, wherein the predictive coding operation comprises sequentially processing each band / bin n of the transform domain representation X(k, n), to determine, for each of the bins, a predicted sample P(k, ri), a prediction residual D(k, ri), and a reconstructed sample X'(k, ri). For each channel k. the processing may proceed in order from high to low frequency bins, or vice versa.
[0079] The predicted sample P(k, ri) is determined by the LC block 206 using linear prediction based on a respective set of reconstructed samples determined for previously processed bins n. In the illustrated example, the LC block 206 implements cross-channel linear prediction, meaning that the LC block 206 determines the predicted sample P(k, ri) using linear prediction based on a respective set of reconstructed samples X'(k, ri) determined for previously processed channels k of a respective set of the other channels of the set of M channels, for a same band n as the respective predicted sample P(k, ri). That is, the LC block 206 performs cross-channel linear prediction based on reconstructed samples determined for previously processed channels for the same bin index n.
[0080] For example, a predicted sample P (k, ri) = PLC(k, ri) may be determined according to:(Eq. 10) where b(n, v) are prediction weights, X'(k — 1 — v, n) is the set of reconstructed samples determined for a number V of other channels of the set of channels (i.e., the number of prediction channels).
[0081] The reconstructed samples X'(k, ri) and prediction residuals D k, ri) may be determined in accordance with Eq. 2 and 3, respectively, as illustrated in Fig. 3. In case the Q and iQ blocks 104, 106 are omitted, the prediction residuals D k, ri) may be determined according to Eq. 3 ’.
[0082] Each reconstructed sample X'(k, n) and prediction residual D(k, n) is stored (e.g., buffered) in the channel buffer 202. Optionally, the channel buffer 202 may also store the transform domain input samples X k, ri).
[0083] Prior to initiating the cross-channel linear prediction, the reconstructed samples X'(k, ni) may be initialized according to Eq. 4 and further according toX'(l,m) = 0 for I < 0(Eq. 11)
[0084] This initialization allows the LC block 206 to use a fixed number of prediction channels V throughout the sequence, although V number of reconstructed samples may not be available until after processing V channels.
[0085] Further, as may be seen from Eq. 10, following this initialization, for the first processed channel k = 0, the predicted samples P(0, n) will become 0, wherein the prediction residuals £)(0, n) become £)(0, n) = Q ~1{Q [2f (0, n)}} (according to Eq. 3) or £)(0, n) = X(0, n) (according to Eq. 3’). Hence, with this initialization, Eq. 11 is defined for each of the channels k = 0... M — 1 and Eq. 10 may thus be applied for processing each of the bins n = 0... N — 1. However, in light of the preceding discussion concerning the initialization of the first channel k = 0, it is noted that the first channel k for which a predicted sample P(k, n) is determined based on a reconstructed sample X'(l, ri) determined for a previously processed bin I = k — 1 as D(l, ri) + P(l, ri), is channel k = 1.
[0086] As an alternative to the initialization given by Eq. 11, the summation range for v in Eq. 10, may be shifted during the course of the sequential processing to include only frequency bands / bins of previously processed channels.
[0087] The set of prediction weights used by the LC block 206 to form a weighted combination of reconstructed samples (e.g., b(n, v) of Eq. 10), may be determined by Coefficient update block 212. For example, denoting like in Eq. 10 a prediction weight for a band n by b(n, v), the set of prediction weights for band n is given by b(n, 0)... b(n, V — 1). The Coefficient update block 212 may determine an updated set of prediction weights for each band n. For example (assuming the prediction process proceeds from high to low frequency bins n), the set of prediction weights b(n — 1,0)... b(n — 1, V — 1) for a band n — 1 may be updated by adjusting each prediction weight of the set of prediction weights b(n, 0)... b(n, V — 1) used for its previously processed neighboring band n by a respective step size. That is, an updated prediction weight for a given band (e.g., b(n — 1, v) for band n — 1) may be determined based on the corresponding prediction weight used for its previously processed neighboring band (e.g., b(n, v for band n) and a respective step size. The set of prediction weights may be stored in an internal buffer of the Coefficient update block 212, or in the channel buffer 202.
[0088] As discussed above with reference to the Coefficient update block 112, a goal of the prediction may be to minimize the energy of the prediction residuals D (k, ri). This may be achieved by updating the set of prediction weights using an LMS method, for instance an NLMS method. More specifically, the respective step size for adjusting each respective prediction weight may be determined using an LMS or NLMS method. For example, the Coefficient update block 212 may determine updated prediction weights according to:b(n — 1, v) = b(n, v) + D (k,ri) g X' (k — 1 — v, n)(Eq. 12) where g is a gain factor and may be given by Eq. 6. The underlying rationale of this approach for updating the prediction weights is the same as discussed above with reference to the Coefficient update block 112.
[0089] Returning to the cross-channel linear prediction and Eq. 10, it is possible to set the number of prediction channels V such that, for each channel, reconstructed samples from all previously processed channels are used for the linear prediction. However, when the number of channels M is large (e.g., M > 30), it may be beneficial (e.g., to reduce computational complexity and memory requirements) to limit the number of prediction channels V to a predetermined smaller number, i.e. V < M. The number of prediction channels V may be determined by channel selector 204. Thus, the channel selector 204 may, when applying a predictive coding operation to a channel k. select a respective set (i.e., subset) of V previously coded channels, for which reconstructed samples X'(k, n) thus are stored in the channel buffer
[0090] For example, upon applying a predictive coding operation to a channel k, the channel selector 204 may determine a set of a predetermined number of the most recently coded channels of the set of channels. As another example, the channel selector 204 may determine an energy for each previously coded channel of the frame, and select as the respective set of previously coded channels, the V the channels having the highest energy. The energy for a channel may for example be computed as e = || " (k, ri) ||2.
[0091] In either case, a respective set of channels selected by the channel selector 204 may include both channels which have been subjected to predictive coding, as well as channels which have been coded using a non-predictive coding operation, such as the first channel k = 0.
[0092] The coding for the set of channels M of step S102 may be concluded after each of the bins N of each of the channels k have been processed. The prediction residuals D(k > 0, n) obtained thereby may represent coded samples for all bins N of all channels k, except for the first channel k = 0, for which the above-mentioned “first reconstructed samples” X'(0, ri) may represent the coded samples. Thus, at step S103 the encoder 200 proceeds to form an encoded bitstream B comprising encoded representations of each of the coded samples of each of the coded channels.
[0093] Analogous to the encoder 100, the encoder 200 may after quantizing the coded samples at Q block 104, apply entropy encoding to each of the (quantized) coded samples "(0, n) and D k > 0, n) by Entropy coding block 116.
[0094] While the entropy coding may provide lossless compression of the coded samples of the coded channels, and hence further may reduce a bit rate of the bitstream B, it is also possible to directly include the (quantized) coded samples into the bitstream B without applying entropy coding, e.g., by multiplexing. The bitstream B may be stored and / or transmitted to a receiving device comprising a decoding system.
[0095] Fig. 4 is a block diagram of a decoding device or system 200’ (hereinafter termed decoder 200’) implementing a method suitable for decoding a bitstream B encoded by the encoder 200, which in the following will be described with further reference to the flow chart of Fig- 9.
[0096] At step S201 the decoder 200’ receives the encoded bitstream B. The decoder 200’ may receive the bitstream B from the encoder 200 via a wired or wireless communication link, or by reading the bitstream B from a memory storing the bitstream B. As explained above with reference to the encoder 200, the bitstream B comprises encoded representations (e.g., quantized and / or entropy encoded representations) of the coded samples generated by the encoder 200 from the transform domain representation for each channel k of the set of channels M of the frame ofthe input signal (e.g., waveform domain input signal). Thus, the coded samples may include the (quantized) coded samples X'(0, ri) and D(k > 0, n) for all bins N.
[0097] At step S202 the decoder 200’ applies at Entropy decoding block 116’ entropy decoding to the encoded bit stream B to obtain the (quantized) coded samples for each channel k, i.e., the (quantized) coded samples X'(0, ri) and D(k > 0, ri) for all bins N. In case no entropy encoding has been applied by the encoder 200, the decoding at step S202 may simply comprise retrieving the coded samples of the respective channels k from the bitstream B, e.g., by demultiplexing.
[0098] Where the coded samples are quantized, the decoder 200’ may as shown include a corresponding inverse quantization iQ block 106 for inverse quantizing the coded samples obtained from the bitstream B.
[0099] At step S203 the decoder 200’ applies a coding operation (which in the decoder also may be referred to as a decoding operation) to each of the channels k, starting with the first channel k = 0. For the first channel k = 0, no prediction or reconstruction is needed, as the for the first channel k = 0, the coded samples are given by the “first reconstructed samples” X'(0,0)... X'(0, N — 1), which simply may be derived from the (entropy decoded) bitstream B and stored in the channel buffer 202.
[0100] After concluding the processing of the first channel k = 0, the decoder 200’ proceeds to process, in sequence, each of the further channels k > 0 using a cross-channel predictive (de)coding operation. The predictive coding operation of the decoder 200’ proceeds analogous to the predictive coding operation of the encoder 200, with the difference that the prediction residuals D (k, ri) here are retrieved from the bitstream B, rather than being derived from predicted samples and input transform domain samples for each channel k as in the encoder 200. Thus, the LC block 206, the Coefficient update block 212, the channel selector 204 and the channel buffer 202 may each operate in the same manner as the correspondingly numbered elements of the encoder 200 in Fig. 3. This further means that the predicted samples P(k, ri) and the reconstructed samples X'(k, ri) also may be determined according to Eq. 10 and Eq. 2.Further, the reconstructed samples X' (k, ni) may be initialized according to Eq. 11. Also, the Coefficient update block 212 may determine updated prediction weights according to Eq. 12 and Eq. 6. The description of the cross-channel predictive coding operation will therefore not be repeated here, but reference is made to the corresponding passages of the preceding description of the encoding method and the encoder 200.
[0101] Hence, for the first channel k = 0 processed by the decoder 200’, “first reconstructed samples” may be determined or initialized to X'(0, n), as retrieved from the bitstream B. For each further channel k > 0, the bins n may be sequentially processed (e.g., in asame order as the encoder 200). Thus, for each bin n of a channel k > 0, a predicted sample P(k, ri) may be determined (e.g., according to Eq. 10) and a reconstructed sample X'(k, ri) may be determined based on the predicted sample P k, ri) and the prediction residual D k, ri). The linear prediction operation proceeds until each of the bins N of the channel k have been processed, wherein a set of N reconstructed samples X'(k, ri) representing the initially encoded channel k of the frame of the input signal have been determined and stored in the channel buffer 202. This set of reconstructed samples X'(k, ri) may be output as a decoded signal of channel k from the decoder 200’.
[0102] In the above, the operations of the encoder 200 and decoder 200’ have been described with reference to a sequential processing of the bands / bins n proceeding in order from high to low frequency bands / bins n, for each channel k. However, it is also possible to sequentially process the bands / bins n for each channel k in order from low to high frequency bands bins n. Since the LC block 206 performs cross-channel prediction, Eq. 10 is applicable also for low to high frequency bin processing (assuming an LC-only coding operation, c.f., the encoder and decoder of Fig. 5 and 6). However, Eq. 12 may assume a modified form:b(n, v) = b(n — 1, v) + D(k, n — 1) g X'(k — 1 — v, n — 1)(Eq. 12’) The prediction weight update of equation Eq. 12’ may equivalently be expressed asbf (n + 1, v) = bf (n, v) + D(k, ri) g X' k — 1 — v,ri)(Eq. 12”) Additionally, the initialization according to Eq. 4 may assume the form of Eq. 4’ (like in the upward linear LP prediction of encoder 100, and decoder 100’).
[0103] Further, in the above exemplified implementations of step SI 02 (performed by the encoder 200) and step S202 (performed by decoder 200’), the encoder / decoder 200 / 200’, for each channel k, and for each bin n, stores the reconstructed samples X'(k, ri) in the channel buffer 202 and determines the prediction residuals D(k, ri) based on the reconstructed samples previously stored in the channel buffer 202. However, in other implementations of step S102 and / or S202, the channel buffer 202 need not store the reconstructed samples X'(k, ri) during processing of each respective bin n, but may instead store the predicted samples P(k, ri) and the prediction residuals D(k, ri) for each bin n. Thus, the LC block 206 may, in analogy with the corresponding discussion of the LP block 110 of the encoder / decoder 100 / 100’, calculate the reconstructed samples X'(k, ri) when needed, e.g., when processing the next band n — 1 (in case of downward prediction) or n + 1 (in case of upward prediction).
[0104] Further, in the above an assumption was made that the transform T used to determine the transform domain representation for each channel k of the M channels has thesame transform length N, such that the number of bands / bins for each channel k is N. However, the cross-channel linear prediction as set out above is also applicable to an implementation wherein different groups of channels are transformed using different transform lengths, e.g., N for a first group ofchannels and N2for a second group of M2channels, etc.. In this case, cross-channel linear prediction may be applied individually to each group of channels (and respective transform blocks within a frame) as set out above, e.g., according to Eq. 10 for each group of channels. This could for instance be useful for applications where different groups of channels of a multi-channel signal may benefit from different transform lengths, such as where some of the channels include transients, while other of the channels are more stationary. In one example, input transform domain samples are segmented into frames of N samples, and N is an integer multiple of both Ah and Ni, such that an integer number of length Ah blocks and an integer number of length N blocks are present in the frame. For instance, for a frame size N equal to 2048, with Ah equal to 256 and Nz equal to 512, there are 8 length Ah blocks and 4 length Nz blocks. Other combinations of frame size N and transform lengths Ah and Nz are possible, and there may be more than two different groups of channels.
[0105] In the above, methods for encoding and decoding using either in-channel (LP) linear prediction at step S102 / S202 or cross-channel (LC) prediction have been disclosed.Methods for encoding and decoding instead using a combination of in-channel (LP) and crosschannel (LC) linear prediction at step S102 / S202 will now be disclosed with further reference to Fig. 5 and Fig. 6, respectively. For conciseness, the terms “LP coding”, “LC coding” and “LP+LC coding” may in the following be used to refer to predictive coding operations including only an in-channel linear prediction, only a cross-channel linear prediction, and a combined in-channel and cross-channel linear prediction, respectively. Hence, the encoder / decoder 100 / 100’ may be referred to as an LP encoder / decoder, and the encoder / decoder 200 / 200’ may be referred to as an LC encoder / decoder.
[0106] Fig. 5 is a block diagram of an encoder 300. The descriptions of the encoder 100 of Fig. 1 and the encoder 200 of Fig. 3 are generally applicable to the encoder 300 of Fig. 5, and thus, like reference signs are used in Fig. 1, Fig. 3 and Fig. 5 to refer to like elements. For conciseness, a description of such like elements will not be repeated, but rather reference is made to the corresponding discussion of Fig. 1 or Fig. 3.
[0107] The encoder 300 however differs from each the encoders 100, 200 in that it includes both an in-channel LP block 110 and a cross-channel LC block 206. The encoder 300 thus implements LP+LC coding and may thus be referred to as an LP+LC encoder.
[0108] The discussion of the input signal in connection with the encoder 200 applies correspondingly to the input signal obtained by the encoder 300, and is thus a multi-channel input signal comprising a plurality of channels M, such as 10 channels or more.
[0109] The encoding method may initially proceed in analogy with the encoding method of the encoder 200. Thus, at step S101 the encoder 300 obtains a respective transform domain representation of each channel k of a frame. Each respective transform domain representation includes a respective set of transform domain samples X(k, ri) for the channel k.
[0110] At step SI 02, the encoder 300 applies a coding operation to each of the M channels. For the first channel k = 0, the encoder 300 may like the encoder 200 apply a non-predictive coding operation. Thus, for the first channel k = 0 the coding operation may simply comprise determining or initializing a set of “first reconstructed samples” X'(0, ri) to correspond to X(0, n), for each bin n. For example, each first reconstructed sample X'(0, ri) may be set to be equal to the input transform domain representation X(0, n), or be equal to Q1[Qf 'CO, n)}}. The thus obtained first reconstructed samples may be stored in the channel buffer 202 as a first set of reconstructed samples for a first coded channel k = 0.
[0111] The further channels k > 0 are processed sequentially, i.e., one after another. For each further channel k > 0 of the frame, a predictive coding operation is applied, wherein the predictive coding operation comprises sequentially processing each band / bin n of the transform domain representation X (k, ri), to determine, for each of the bins, a predicted sample P (k, ri), a prediction residual D(k, ri), and a reconstructed sample X'(k, ri). The processing may proceed in order from high to low frequency bins, or vice versa.
[0112] As illustrated in Fig. 5, each predicted sample P(k, ri) may in the encoder 200 be determined based on both an in-channel linear prediction and a cross-channel linear prediction. For instance, a predicted sample P(k, ri) may be determined according to:P(k,ri) = PLPk,ri) + PLC(k,n)(Eq. 13) where PLP(k, ri) represents the in-channel linear prediction provided by the LP block 110 and PLC(k,n) represents the cross-channel linear prediction provided by the EC block 206. The sum of these predictions, P(k, ri), is determined by the sum block 302. In case of downward linear prediction, PLP(k, ri) may be determined according to Eq. 1 and PLC(k, ri) may be determined according to Eq. 10. Accordingly, a predicted sample P(k, ri) may be determined according to:(Eq. 14)
[0113] Each reconstructed sample X'(k, ri) is stored (e.g., buffered) in the sample buffer 114. Each reconstructed sample X'(k, ri) is further stored in the channel buffer 202 to facilitate the cross-channel linear prediction. Further, in the illustrated example, each prediction residual D (k, ri) is stored in the channel buffer 202 but may alternatively or additionally be stored in the sample buffer 114. While in the block diagram of Fig. 5, the sample buffer 114 and the channel buffer 202 are shown as separate blocks, this is merely an example, and it is to be understood that their respective functionalities may be implemented in a common frame buffer. This applies also to the block diagram of Fig. 6 showing the decoder 300’, discussed below.
[0114] Prior to initiating the predictive coding operation for a channel k > 0, the reconstructed samples X'(k, m) may be initialized according to Eq. 4 and Eq. 11.
[0115] The prediction coefficients a(n, u) and b(n, v) may be determined and updated by Coefficient update block 312. For example, the set of prediction coefficients a(n, u) used by the LP block 110 may be determined according to Eq. 5 while the set of prediction coefficients b(n, v) used by the LC block 206 may be determined according to Eq. 12. However, the gain factor g may here be determined according to Eq. 6, where the norm || " (k, n) || is determined according to:|| (k,n)|| =k(Eq. 15)
[0116] Corresponding expressions may be derived for the case where upward prediction is used, i.e., where the prediction of the LP block 110 and the LC block 206 proceeds from low to high frequency bins.
[0117] When applying a predictive coding operation to a channel k, the channel selector 204 may, as described with reference to the encoder 200, be used to select a respective set (i.e., subset) of V previously coded channels, for which reconstructed samples X'(k, ri) thus are stored in the channel buffer 202, based on which the LC block 206 may determine the cross-channel linear prediction PLC(k, ri).
[0118] The coding for the set of channels M of step S102 may be concluded after each of the bins N of each of the channels k have been processed. The prediction residuals D(k > 0, n) obtained thereby may represent coded samples for all bins N of all channels k, except for the first channel k = 0, for which the above-mentioned “first reconstructed samples” X'(0, ri) may represent the coded samples. Thus, at step S103 the encoder 300 proceeds to form an encoded bitstream B comprising encoded representations of each of the coded samples of each of the coded channels.
[0119] Analogous to the encoders 100, 200, the encoder 300 may after quantizing the coded samples at Q block 104, apply entropy encoding to each of the (quantized) coded samples by Entropy coding block 116. However, it is also possible to directly include the (quantized) coded samples into the bitstream B without applying entropy coding, e.g., by multiplexing. The bitstream B may be stored and / or transmitted to a receiving device comprising a decoding system.
[0120] In the above, the first channel k = 0 was coded using a non-predictive coding to initialize the “first reconstructed samples” X'(0, ri). However, the presence of the LP block 110 enables another initialization approach for the first channel k = 0, namely the initialization approach set out above with reference to the LP block 110. Thus, for the first channel k = 0, a “first reconstructed sample” X'(k, n0), or a “first set of L reconstructed samples”, may be set to be equal to the input transform domain representation X(k, n0), or be equal toQ-1{Q{X(k, n0)}}, where n0denotes the highest or lowest frequency bin or bins. For example, for downward prediction, n0= n — 1 and for upward prediction n0= 0, assuming L = 1.Thereafter, the further bins n of the first channel k = 0 may be processed using the in-channel linear prediction of the LP block 110, e.g., according to Eq. 1 (corresponding to the first sum of Eq. 14). Thus, after concluding the coding of the first channel k = 0, coded samples including the prediction residuals £)(0, ri) have been obtained using Eq. 1, for all bins N of the channel k, except for the L bin(s) n0, for which the respective coded sample is given by X'(0, n0).
[0121] Fig. 6 is a block diagram of a decoding device or system 300’ (hereinafter termed decoder 300’) implementing a method suitable for decoding a bitstream B encoded by the encoder 300, which in the following will be described with further reference to the flow chart of Fig. 9. The decoder 300’ implements LP+LC coding and may thus be referred to as an LP+LC decoder.
[0122] The decoding method may initially proceed in analogy with the decoding method of the decoder 200’. Thus, at step S101 the decoder 300’ receives the bit stream B from the encoder 300 via a wired or wireless communication link, or by reading the bitstream B from a memory storing the bitstream B. As explained above with reference to the encoder 300, the bitstream B comprises encoded representations (e.g., quantized and / or entropy encoded representations) of the coded samples generated by the encoder 300 from the transform domain representation for each channel k of the set of channels M of the frame of the input signal (e.g., waveform domain input signal).
[0123] At step S202 the decoder 300’ applies at Entropy decoding block 116’ entropy decoding to the encoded bit stream B to obtain the (quantized) coded samples for each channel k. for all bins N. In case no entropy encoding has been applied by the encoder 300, the decoding atstep S202 may simply comprise retrieving the coded samples of the respective channels k from the bitstream B, e.g., by demultiplexing.
[0124] Where the coded samples are quantized, the decoder 300’ may as shown include a corresponding inverse quantization iQ block 106 for inverse quantizing the coded samples obtained from the bitstream B.
[0125] At step S203 the decoder 300’ applies a coding operation (which in the decoder also may be referred to as a decoding operation) to each of the channels k. starting with the first channel k = 0. For the first channel k = 0, no prediction or reconstruction is needed in case the first channel k = 0 has been coded using non-predictive coding. Thus, the coded samples for the first channel k = 0 may be given by the “first reconstructed samples” X'(0,0)... X'(0, N — 1), which simply may be derived from the (entropy decoded) bitstream B and stored in the channel buffer 202 as a first set of reconstructed samples for the first coded channel k = 0.
[0126] Alternatively, if for the first channel k = 0, the initialization approach set out above with reference to the LP block 110 was used, for the first channel k = 0, a “first reconstructed sample” X'(k, n0), or a “first set of L reconstructed samples”, may correspond to the input transform domain representation X(k, n0), (e.g., equal to X(k, n0) or Q-1{Q{X(k, n0)}}). Thus, no reconstruction is needed for band or bands n0for channel k = 0 as the first reconstructed sample(s) X'(k, n0) (e.g., the first set of L reconstructed samples) simply may be retrieved from the bitstream B. Thereafter, the further bins n of the first channel k = 0 may be processed using the in-channel linear prediction of the LP block 110, e.g., according to Eq. 1 (corresponding to the first summation of Eq. 14). Thus, after concluding the (de)coding of the first channel k = 0, coded samples including the reconstructed samples X'(0, ri) have been obtained based on the predicted samples P(k, n) and corresponding prediction residuals D(k, n0), for all bins N of the channel k, except for the bin(s) n0, for which the coded sample(s) is given by X'(0, n0) as retrieved from the bit stream B. The thus obtained reconstructed samples for the first channel k = 0 may be stored in the channel buffer 202 as a first set of reconstructed samples for the first coded channel k = 0.
[0127] After concluding the processing of the first channel k = 0, the decoder 300’ may proceed to process, in sequence, each of the further channels k > 0 using both in-channel and cross-channel predictive (de)coding operation. The predictive coding operation of the decoder 300’ proceeds analogous to the predictive coding operation of the encoder 300, with the difference that the prediction residuals D (k, ri) here are retrieved from the bitstream B, rather than being derived from their corresponding predicted samples and input transform domain samples for each channel k, as in the encoder 300. Thus, the LC block 206, the Coefficient update block 312, the channel selector 204 and the channel buffer 202 may each operate in thesame manner as the correspondingly numbered elements of the encoder 300 in Fig. 5. This further means that the predicted samples P k, ri) and the reconstructed samples X'(k, ri) also may be determined according to Eq. 13 (or 14) and Eq. 2, respectively. Further, the reconstructed samples X'(k, ni) may be initialized according to Eq. 4 and Eq. 11. Also, the Coefficient update block 212 may determine updated prediction weights according to Eq. 12 and Eq. 6. The description of the combined in-channel and cross-channel predictive coding operation will therefore not be repeated here, but reference is made to the corresponding passages of the preceding description of the encoding method and the encoder 300.
[0128] Just like it is possible to apply only an in-channel predictive coding operation to the first channel k = 0 in the encoder 300 and decoder 300’, it is also possible to vary the type of coding operation between the further channels. For instance, for each channel k > 0 the encoder may determine on a channel-by-channel basis whether to apply LP coding, LC coding or both (LP+LC coding). This may be done within a single frame. The encoder 300 may, for example, during operation determine the type of coding to apply for a frame based on evaluating a respective prediction gain determined for each channel k > 0 based on the set of input transform domain samples X(k, ri) and the prediction residuals D(k, ri) determined by processing the channel k using a given coding approach, e.g., LP coding, LC coding or LP+LC coding. While it in principle is possible to determine a respective prediction gain for each channel k for each of the three coding approaches, this may be computationally too expensive in some applications. Hence, a more effective approach may be to determine a respective prediction gain pg(k) for each channel k using, for example, only LP+LC coding, and only if a sufficient prediction gain (i.e., the prediction gain pg(k) exceeding a threshold Tgsuch as 1.0) is achieved, may the LP+LC coded samples be included in the bit stream. In case a sufficient prediction gain is not obtained for the LP+LC coding, the encoder 300 may determine to use (only) LC coding. LC coding tends to demand less computational resources than LP coding (since the number of prediction channels V typically is smaller than the prediction order U). Hence, it may be useful to apply LC coding to the channel k even if LP+LC does not result in a sufficient prediction gain pg(k). The LC coding of the channel k may performed responsive to determining that LP+LC coding does not provide a sufficient prediction gain. However, it is also possible to perform both LC and LP+LC coding and determine, based on the evaluation of the prediction gain pg(k), whether to include either the prediction residuals D (k, ri) determined using LC coding, or the prediction residuals D (k, ri) determined using LP+LC coding in the bit stream B for the channel k.
[0129] Optionally, the encoder may, based on the evaluation of the prediction gain pg(k) for a channel k, determine to not include any predictive coded samples (e.g., prediction residuals D(k, n)) in the bitstream B at all, but rather include the input transform domain samples X(k, ri) or Q{X(k, n)} for the channel k in the bitstream B. A similar approach may be applied to an LC encoder (e.g., encoder 200), as well as to an LP encoder (e.g., encoder 100) provided the input signal is a multi-channel signal. In this case, the encoder 100 or 200 may determine a prediction gain for each channel k and include LP or LC coded samples for the channel k in the bitstream B only if a sufficient prediction gain is obtained, and otherwise include the input transform domain samples X (k, ri) or Q{X (k, n)} for the channel k in the bitstream B.
[0130] A prediction gain pg(k) for a channel k may for instance be determined based on a ratio of ||X(k, n)|| and ||D(k, n)||. For example, the prediction gain pg(k) may be determined according to pg(k) = ||X(k, n)||2 / ||D(k, n)||2. When determining the prediction gain pg(k) it may be beneficial to use Eq. 3’ to derive D(k, n), to avoid that any loss introduced by the dequantization and quantization operations reduces the reliability of the prediction gain pg(k). Thus, this applies also to a case where the encoder 300 includes Q block 104 and iQ block 106. Within the context of the block diagram of Fig. 5, this may correspond to obtaining ||X(k, n)|| — ||D(k, n)|| directly from the output from the difference block 102 (e.g., and store it in the sample buffer 114, for instance). It is to be noted that the prediction gain pg(k) for a channel k need not be based on ||X(k, n) || and \\D (k, ri) || for each of the bands n, but also may be determined for a subset of the bands n, such as a subset of the bands n which are assumed to have the highest energy among the bands n for the channel k.
[0131] A prediction gain pg(k) and a threshold comparison as presented above may also be used to determine what type of coding to apply for a single-channel input signal (e.g., in LP encoder 100) or all channels of a multi-channel signal (e.g., in LC encoder 200 or LP+LC encoder 300). That is, upon obtaining a current frame, the encoder (e.g., encoder 100, 200 or 300) may determine a prediction gain pg(k) for the single channel, or for a given channel of a multi-channel signal (e.g., a first channel k encoded using LP, LC, or LP+LC encoding). The encoder may further evaluate the prediction gain pg(k) by comparing the prediction gain pg(k) to a threshold Tg(e.g., Tg= 1.0) to determine whether the encoded bit stream, for the current frame, is to be formed to comprise coded samples in the form of the prediction residuals (determined by LP, LC or LP+LC coding) or non-prediction coded samples, i.e., the input transform domain samples X(k, ri) or Q{X(k, n)} for each channel k (or the single channel).
[0132] Hence, to sum up, with reference to an encoder (LP, LC or LP+LC encoder) receiving a plurality of channels k of a frame of a multi-channel signal, the encoding method may comprise:for each respective channel k of at least a subset of channels of the plurality of channels: applying a predictive coding operation to the respective channel k to determine coded samples in the form of prediction residuals D(k, ri) for the channel k, andforming an encoded bitstream comprising the coded samples for each channel of the at least a subset of channels.
[0133] In case the encoder is an LP encoder, LP encoding may be applied to each channel of the at least a subset of channels. In case the encoder is an LC encoder, LC encoding may be applied to each channel of the at least a subset of channels. In case the encoder is an LP+LC encoder, any one of LP encoding, LC encoding or LP+LC encoding may be applied to each channel of the at least a subset of channels. Each channel may be encoded using a same type of predictive coding operation (i.e., LP, LC or LP+LC). The type of predictive coding operation may also be determined on a per-channel basis. In either case, a determination of the type of predictive coding operation (i.e., LP, LC or LP+LC) may be determined based on a prediction gain. Where a same type of predictive coding operation is applied to each channel, the prediction gain may be determined for a given channel of the at least a subset of channels. Where the type of predictive coding operation is determined on a per-channel basis, a respective prediction gain may be determined for each respective channel (e.g., based on the input transform domain samples for the channel and the prediction residuals for the channel, as discussed above). In either case, the prediction gain may be compared to a prediction gain threshold and the type of predictive coding operation may be determined based on whether the (respective) prediction gain exceeds the prediction gain threshold or not.
[0134] In some examples, the at least a subset of channels may be a first subset of the plurality of channels and the plurality of channels may further comprise a second subset of channels wherein the encoding method may further comprise, for each respective channel k of the second subset of channels determine coded samples in the form of representations of transform domain samples of the respective channel (e.g., the representations being the input transform domain samples or quantized representations of the input transform domain samples X(k, n)), wherein the encoded bitstream further is formed to comprise the coded samples for each channel of the second subset of channels.
[0135] The second subset of channels may be channels for which a prediction gain is less than a threshold. Accordingly, the encoding method may comprise, to each channel of the plurality of channels, applying a predictive coding operation to the respective channel k todetermine prediction residuals D(k, ri) for the channel k, determining a prediction gain for the respective channel, and identifying the channel as a channel of the first subset of channels responsive to the prediction gain exceeding a threshold and otherwise identifying the channel as a channel of the second subset of channels. In this case, and further assuming that the encoder is an LC or LP+LC encoder, the linear prediction operation for at least some of the first subset of channels may be further based on coded samples for previously processed bands of the second subset of channels. That is, a cross-channel linear prediction of a predicted sample P(k, n) for a channel k of the first subset of channels may be determined using linear prediction based on a respective set of reconstructed samples X'(k, ri) determined for a respective subset of previously processed channels k of the first subset of channels, and coded samples determined for a respective subset of previously processed channels of the second subset of channels, for a same band n as the respective predicted sample P(k, ri).
[0136] As may be appreciated from the preceding discussion, regardless of whether the encoder, for a current frame, determines to encode prediction residuals into the bit stream B or not, the predictive coding operation (e.g., LP, LC and / or LP+LC, depending on the type of encoder) may still be performed for the current frame, at least for one channel k.
[0137] The encoder may indicate to the decoder whether the channel has been subjected to linear predictive coding or not by including in the bit stream B an indicator indicating whether the channel has been subjected to linear predictive coding or not. The indicator may be an explicit flag, e.g., a single bit flag or a multi-bit flag. In the case of an LP+LC encoder, a multibit flag (e.g. 2-bits) may be used to further indicate a type of predictive coding (LP, LC or LP+LC). The flag may be included on a per-frame basis, on a per-transform block basis, or on a per-channel basis if the encoder has the ability to vary the type of coding operation between the channels.
[0138] Fig. 7a is a high-level block diagram of an encoder system 400 suitable for encoding a multi-channel waveform domain input signal W, such as a time domain input signal. The encoder system 400 comprises a transform stage 402 for transforming each channel k of the waveform domain input signal W to obtain, per frame, respective transform domain representations for each of the channels k. In the illustrated example, the transform stage 402 is realized using a set of DCT-blocks, however this is merely an example and any of the other types of transforms discussed above may be used instead. Fig. 7a shows three channels however this is merely an example.
[0139] The transform domain representations of each channel k are provided to Prediction & Quantization block 404, which may be realized in accordance with any one of the encoders of Fig. 1, Fig. 3 or Fig. 5.
[0140] The Prediction & Quantization block 404 outputs coded samples for each channel k to Entropy coding block 406, which entropy encodes the coded samples for each channel k into an entropy encoded bitstream B.
[0141] The encoder system 400 may as shown further comprise a Bitstream multiplexer (muxing) block 408 multiplexing an indicator such as a flag (possibly a multi-bit flag) into the entropy encoded bitstream B indicating the type of coding (e.g., LP, LC or LP+LC) applied by the Prediction & Quantization block 404, either on a per-channel basis, or jointly for all channels. In either case, the flag(s) may be provided on a per-frame basis, or on a per-transform block basis. Thus, the type of coding applied by the Prediction & Quantization block 404 may vary on a frame-by-frame basis, and be signaled to the decoder.
[0142] Fig. 7b shows a high-level block diagram of a corresponding decoder system 500. The encoder system 400 and the decoder system 500 may together define an encoder-decoder system. The decoder system 500 mirrors the encoder system 400 and thus comprises a bitstream demultiplexing (demuxing) block 502 for demultiplexing any indicator(s) (e.g., flag(s)) added to the encoded bitstream B by the encoder 400. The coded samples for each channel k are provided to the Entropy decoding block 504, decoding the entropy encoded bitstream B into a respective set of coded samples for each channel k.
[0143] Each set of coded samples for each channel k is in turn provided to Prediction & Quantization block 506, which may be realized in accordance with any one of the decoders of Fig. 2, Fig. 4 or Fig. 6.
[0144] The Prediction & Quantization block 506 outputs decoded (e.g., reconstructed) samples for each channel k which in turn may be provided to inverse transform stage 508, performing an inverse transform (e.g., iDCT) to transform the decoded samples for each channel k into waveform (e.g., time) domain samples for each channel k, representing a reconstruction of the original input signal W input to the encoder 400.
[0145] The person skilled in the art realizes that the present disclosure by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims.
[0146] For example, as mentioned above, the encoders and decoders may operate on a frame-by-frame basis. Hence, reference has in the above mainly been made to a single frame. However, it is to be noted that during operation an encoder or decoder may typically receive a sequence of frames, to be sequentially processed by the encoder or decoder. As the encoding and decoding methods do not rely on information derived from previous frames, the predictor states (e.g., the sets of prediction weights, the reconstructed samples, etc.) may be initialized (e.g., reset) to zero between successively received frames.
[0147] For example, while in the discussion of the encoders and decoders of Fig. 1-6, the predictive coding operations have been disclosed to be applied to all bins N, a predictive coding operation may in some implementations be selectively applied to one or more frequency ranges of the transform domain representation of a channel k of an input signal. That is, the set of all bands / bins N of the transform domain representation may be partitioned into a number (e.g., two or more) of subsets of (consecutive) bands / bins spanning a respective predetermined frequency range. In this case, analogous to the above discussion, LP, LC or LP+LC coding may be applied on a per frequency range basis, e.g., selectively. Of course, in case of the LP encoder / decoder 100 / 100’ only LP coding or non-predictive coding may be applied to a frequency range.Correspondingly, in case of the LC encoder / decoder 200 / 200’ only LC coding or non-predictive coding may be applied to a frequency range. For instance, an encoder may apply predictive coding (LP, LC or LP+LC coding) only to one or more frequency ranges for which a prediction gain is achieved. Whether to apply predictive coding, and / or the type of predictive coding to apply, to a given frequency range may be indicated in the bitstream B by the encoder including an indicator on a per- frame, per-block, per-channel, and / or per- frequency range basis (analogous to the discussion of indicators and flags above). It is to be noted that upon initiating a coding operation for each frequency range, the predictor states (e.g., the sets of prediction weights, the reconstructed samples, etc.) may be initialized (e.g., reset) to zero. The above description of LP coding, LC coding, and LP+LC coding with reference to Fig. 1-6 applies correspondingly to each subset of bins I frequency range for which either LP, LC or LP+LC coding is applied. Also, the discussion of upward and downward linear prediction is applicable to each such subset of bands, wherein the highest or lowest frequency bin (or bins) n0may refer to the highest of lowest frequency bin (or bins) of respective subset of bands.
[0148] As discussed above, e.g., with reference to the Coefficient update blocks 112, 212 and 312, a predictive coding operation may include determining an updated set of prediction weights for each respective band. In some implementations, updating the set of prediction weights for a respective band may comprise adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, by a respective step size. Thus, an updated prediction weight of an updated set of prediction weights for a respective band may be determined based on a prediction weight (“previous prediction weight”) of a set of prediction weights (“previous set of prediction weights”) used for a previously processed band neighboring to the respective band, and a respective step size. The updated prediction weight may be determined to be equal to the previous prediction weight adjusted by the step size. In other words, the updated prediction weight may be determined as thesum of the previous prediction weight and the respective size. This is exemplified by, for instance, Eqs. 5, 5’, 12, 12’ and 12”.
[0149] The updated prediction weight and the previous prediction weight in the above may in particular refer to corresponding prediction weights of the updated and the previous sets of prediction weights, respectively.
[0150] As further exemplified by Eqs. 5, 5’, 12, 12’ and 12”, each prediction weight of a set of prediction weights may be associated with a respective index (e.g., u in Eqs. 5 and 5’, or v in Eqs. 12, 12’ and 12”). The index may also be referred to as a prediction weight index or a prediction coefficient index (typically corresponding to the so-called “tap index” of the given LP or LC predictor). As used herein, the term “corresponding prediction weights” may be understood as prediction weights associated with the same index. Accordingly, the updated prediction weight and the previous prediction weight in the above may in particular refer to prediction weights of the updated and the previous sets of prediction weights associated with the same index (e.g., the same index u or the same index v).
[0151] As may be appreciated from the above, the number of prediction weights in a given set of prediction weights corresponds to the prediction order (e.g., U or 7). In some instances, the prediction order may be 1 (corresponding to a first-order predictor), wherein the set of prediction weights may be defined by a single prediction weight (e.g., a singleton set of prediction weights). This applies correspondingly to a set of previously coded samples used for determining a predicted sample for a given band. Thus, in case the prediction order is one, the set of previously coded samples may be defined by a single previously coded sample (e.g., a singleton set of previously coded samples).
[0152] By adaptively determining the prediction weights as set out herein in the same manner during both encoding and decoding, the prediction weights need not be encoded into the bitstream B.
[0153] As mentioned above, the encoding and decoding methods discussed above do not rely on information derived from previous frames, and hence the predictor states (e.g., the sets of prediction weights, the reconstructed samples, etc.) may be initialized (e.g., reset) to zero between successively received frames. However, in some implementations, information from a preceding frame may be used for initializing the prediction weights during encoding and decoding, as set out in the following:
[0154] For ease of explanation, reference is in the following made to a first frame f=fi and a second frame f=f2, wherein the first and second frames f1 and f2 may be consecutive frames of a sequence of frames.
[0155] The prediction weight initialization approaches set out in the following may be used in encoding and decoding implementations using (at least) linear cross-channel prediction (LC prediction). Thus, it is in the following assumed that each frame of the input signal (as processed during encoding) and the encoded bitstream (as processed during decoding) includes a set of channels. The following disclosure may, for example, be implemented in the encoders and decoders of Figs. 3-6. More specifically, the initialization of the prediction weights may be implemented by the respective coefficient update blocks 212 and 312 shown therein. The set of channels may comprise a relatively large number of channels (e.g., 10 or more, 30 or more, 100 or more), for instance, in biomedical signal applications. However, the following disclosure is applicable also to a smaller number of channels, such as 3 or more channels.
[0156] The encoding may proceed as discussed above with reference to the encoders 200 and 300. Thus, for each of the frames / ; and / 2, a transform domain representation of the frame / is obtained that includes, for each channel k, of the set of channels, a set of transform domain samples (that interchangeably may be referred to as transform domain coefficients). A transform domain sample / coefficient associated with a frame / a channel k, and a band (or bin) n is in the following denoted Xf (k, n).
[0157] A predictive coding operation is then applied to each given channel k of the frame / that is to be encoded using LC or LP+LC prediction, to determine a set of predicted samples, a set of prediction residuals, and a set of reconstructed samples for each channel. Analogous to the transform domain coefficients Xf (k, n), a predicted sample, a prediction residual and a reconstructed sample associated with a frame / a channel k. and a band n is in the following denoted Pf(k, n), D (k, n) and X' (k, n), respectively. The predictive coding operation comprises sequentially processing the bands for the given channel k in a direction from an initial band, n o, to a final band, n i.
[0158] Thereby, for frame / channel k and band n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples (one or more) determined for the given band n, for a set of previously coded channels (one or more) of the frame / ;a prediction residual, D (k, n), is determined based on the transform domain coefficient Xf (k, n) and the predicted sample Pf(k, n) determined for the given band, n, and represents a prediction error between the transform domain coefficient / (k, n) and the predicted sample P (k, n); anda reconstructed sample, X'f (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) determined for the given band n.
[0159] In case only LC prediction is used, a predicted sample Pf(k, ri) may be given by Pf(k, ri) = PfLC(k, n). IncaseLP+LC prediction is used, a predicted sample Pf(k, ri) may be given by Pf(k, ri) = PfLP(k, ri) + PfLC(k, ri). An LC prediction PfLC(k, ri) may be determined according to Eq. 10. An LP prediction PfLP(k, ri) may be determined according to Eq. 1 or 1’.
[0160] Further, and still with reference to the frame / , and channel k, for each band n except the initial band nf,0, an updated set of prediction weights is determined for the given band n. As discussed above, each prediction weight of the updated set of prediction weights may be determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size. It is noted that, for a downward prediction direction (meaning that n o > n i), m = n + 1. For an upward prediction direction (meaning that n o < n i), m = n — 1. In case of a downward prediction direction, the prediction weights may be updated in accordance with Eq. 12. In case of an upward prediction direction, the prediction weights may be updated in accordance with Eq.12’ or (equivalently) 12”. In case of an LP+LC prediction, the prediction weights for the LP prediction may be updated in accordance with Eq. 5 (in case of a downward prediction direction) or Eq. 5’ (in case of an upward prediction direction).
[0161] For the initial band np.o in the second frame f2, the set of prediction weights may be initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band nf1,1for the given channel k, in the first frame f1, wherein the final processed band nf1,1in the first frame f1 is the same band as the initial processed band nf2,0in the second frame f2;iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f1.
[0162] For each of the first and second frames / ? and / 2, the prediction errors (typically the quantized representations Q{Xf(k, n) — Pf(k, n)} determined for each band n and each channel k are encoded into the bitstream B.
[0163] Decoding may proceed in a corresponding manner, as discussed above with reference to the decoders 200’ and 300’. Thus, an encoded bitstream to be decoded is received that comprises, for each of the frames / ? and / 2, encoded representations of a set of prediction residuals for each channel k.
[0164] The encoded bitstream B is decoded (e.g., entropy decoded) to obtain the set of prediction residuals for each channel k wherein each prediction residual Df (k, n) is associated with a respective band n. As explained above, where, during encoding, each prediction residual Df(k, n) is determined according to Df(k, ri) = Q-1{Q(X(k, n) — Pf(k, n))}, the prediction residuals may correspondingly be obtained by dequantizing the quantized prediction errors Q{X(k, n) — P(k, ri)} decoded from the bitstream (e.g., by the dequantization block 106 of Fig.4 or Fig. 6).
[0165] A predictive coding operation is then applied to each given channel k of the frame / that is to be decoded using LC or LP+LC prediction (in accordance with the type of prediction used during encoding), to determine a set of predicted samples and a set of reconstructed samples for each channel. The predictive coding operation comprises sequentially processing the bands for the given channel k in a direction from an initial band, n o, to a final band, n i.
[0166] Thereby, for frame / channel k and band n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples (one or more) determined for the given band n, for a set of previously coded channels (one or more) of the frame / ; anda reconstructed sample, X' (k, n), is determined based on the predicted sample Pf(k, ri) and the prediction residual Df (k, ri) associated with the given band n and decoded from the bitstream.
[0167] Further, the set of prediction weights may be initialized and updated in the same manner as described above in relation to the encoding method. The set of reconstructed samples X'(k, n) may be output as a decoded signal for channel k.
[0168] Initialization options i, ii, iii of the encoding and decoding methods set out above, may in the following be termed “mode 1”, “mode 2” and “mode 3”, respectively. It is noted that, “the set of prediction weights” discussed in the following refers to the set of prediction weights used for the LC prediction, unless stated otherwise.
[0169] Mode 1: According to this initialization mode, for each given channel k to be coded using LC prediction within the second frame / 2, each prediction weight of the set of prediction weights for the initial band nf2,0is initialized with a zero weight. In case the set of prediction weights is given by Eq. 12, 12’ or 12”, mode 1 involves setting:(Eq. 16)for each prediction weight index v. Mode 1 does not rely on information from the preceding frame f1. Mode 1 is hence an initialization mode that may be used in any instance where information from a preceding frame is not available (e.g., since LC prediction was not used in the preceding frame), or where resetting of the predictor states is desired. Of course, initialization mode 1 may be used also for the temporally first frame in a sequence of frames to be coded using LC prediction. A generalized form of the initialization according to Eq. 16 that is defined for any frame f is given by:bf,k(nf,0,v) = 0(Eq. 16’)
[0170] Mode 2: According to this initialization mode, for each given channel k to be coded using LC prediction within the second frame / 2, each prediction weight of the set of prediction weights for the initial band np.o is initialized with the value of its corresponding prediction weight (i.e., associated with the same prediction weight index v) determined for the final band nfi,i in the first frame / ;. In case the set of prediction weights is given by Eq. 12, 12’ or 12”, mode 2 involves setting:bf2,k(nf2,0, v) = bf1,k(nf1,1, v)(Eq. 17) for each prediction weight index v.
[0171] Mode 2 hence utilizes information about the updated set of prediction weights determined for the final band of the preceding frame for initializing the predictor states for the current frame. Mode 2 may be utilized where the LC prediction alternates between upward prediction (i.e., from low to high frequency bands) and downward prediction (i.e., from high to low frequency bands) between the first and second frames / ; and / 2, and where the final processed band n / ;,; in the first frame / ; is the same band as the initial processed band n / 2,oin the second frame f. (i.e., n / ;,; = n / 2,0). Typically, this also means that the prediction order for the LC prediction (e.g., V in Eqs. 12, 12’ and 12”) is the same in the first and second frames f1 and f2. Hence, the initial processed band nf1,0 in the first frame f1 may be the same band as the final processed band nf2,1 in the second frame f2. (i.e., nf1,0 = nf2,1).
[0172] In some implementations, this initialization approach may be extended across multiple frames. Suppose the first and second frames / ;, / 2 are comprised in an alternating sequence of first frames and second frames / . For each channel k that is to be coded using LC prediction, the bands n are sequentially processed in a direction from an initial band, n o, to a final band, nfi. The prediction direction alternates between upward and downward prediction such that in each first frame nfo < n i (meaning upward prediction) and in each second frame n o > nfi (meaning downward prediction), or vice versa. Suppose further that for each frame / theinitial band nf,0 is the same band as the final band nf-1,1 of its preceding frame f-1. Using initialization mode 2, the set of prediction weights for the initial band nf,0, (for each given channel k) in a frame f may thus be initialized with the set of prediction weights used for processing the final band nf-1,1 for the given channel k, in the preceding frame f-1. In case the set of prediction weights is given by Eq. 12, 12’ or 12”, mode 2 may be expressed by the following generalized form of Eq. 17:(Eq. 17’) for each prediction weight index v.
[0173] As may be appreciated from the above, using alternating prediction directions, in combination with the initialization according to mode 2, allows the values of the finally updated states of the set of prediction weights for each channel k in a preceding first frame f-1 to be used as a starting point for the set of prediction weights used for coding the channel k in the next second frame f. It is contemplated that, for many types of signals (in both biomedical applications and audio applications), some degree of signal correlations may typically exist across frame boundaries within a given channel. At least, it may be expected that some signal characteristics are shared across frame boundaries. Therefore, it is envisaged that the updated prediction weights used to code the final band nf-1,1, in frame f-1 may represent a better starting point for the prediction weights used to code the initial band nf,0 in frame f and thus provide an improved coding gain already for the first few coded bands in frame f compared to initializing the prediction weights with zero-valued weights according to mode 1.
[0174] Mode 3: According to this initialization mode, for each given channel k to be coded using LC prediction within the second frame f2, each given prediction weight of the set of prediction weights for the initial band nf2,0 is initialized with a respective average prediction weight value computed as an average of at least a subset of its corresponding prediction weights (i.e., associated with same prediction weight index v) determined for the channel k in the first frame f1. The at least a subset of corresponding prediction weights may here refer to the corresponding prediction weights (i.e., associated with the prediction weight index v) determined for at least a subset of the bands from the initial band nf1,0 to the final band nf1,1 in the first frame f1. In case the set of prediction weights is given by Eq. 12, 12’ or 12”, mode 3 involves setting:(Eq. 18)for each prediction weight index v, where PI is at least a subset of the bands from the initial band nfi,o to the final band nf.i,i in the first frame; and Ntotis the number of bands of the at least a subset of bands. In some implementations, the average prediction weight value is computed as an average of each of its corresponding prediction weights determined for the channel k in the first frame / ;, i.e., the corresponding prediction weights determined for each band from the initial band ni,o to the final bandIn this case, Eq. 18 may be re-written as:(Eq. 18’) and further assuming that
[0175] In some implementations, this initialization approach may be extended across multiple frames. Suppose the first and second frames f1, f2 are comprised in a sequence of frames f. For each channel k that is to be coded using LC prediction, the bands n are sequentially processed in a direction from an initial band, nf,0, to a final band, nf,1. The prediction direction may be fixed or vary between frames. Using initialization mode 3, each prediction weight of the set of prediction weights for the initial band nf,0, (for each given channel k) in a frame f may thus be initialized with a respective average prediction weight value computed based on its corresponding prediction weights in the preceding frame f-1. In case the set of prediction weights is given by Eq. 12, 12’ or 12”, mode 3 may be expressed by the following generalized form of Eq. 18:(Eq. 18”)
[0176] As may be appreciated from the above, mode 3, like mode 2, hence utilizes information about the updated set of prediction weights determined for the final band of the preceding frame for initializing the predictor states for the current frame. On a similar rationale as presented with reference to mode 2, it is envisaged that the average prediction weight values derived from the prediction weights used to code the bands PI in frame f-1 may represent a better starting point for the prediction weights used to code the initial band n / o in frame / and thus provide an improved coding gain already for the first few coded bands in frame / compared to initializing the prediction weights with zero-valued weights according to mode 1.
[0177] Whereas mode 2 pre-supposes an alternating prediction direction, mode 3 may be used with either fixed or alternating prediction directions. Therefore, mode 3 may advantageously be used in combination with a downward prediction direction (e.g., meaning thatTif ^ < n^0for at least frames / ; and / 2). The benefits of mode 3 may thus be combined with the typically improved prediction gain a downward prediction direction confers over an upward prediction direction. A further benefit of mode 3 is that it may be used also where the prediction order varies between frames.
[0178] To enable using initialization mode 2 or 3 in a next frame the sets of prediction weights determined for each band n in a given frame f-1 may be stored in a buffer such that the prediction weights needed for to perform the initialization can be retrieved upon processing the next frame / . During encoding, the prediction weights may for example be stored in an internal buffer of the Coefficient update block 212 of Fig. 3 or block 312 of Fig. 5, or in the channel buffer 202 of Fig. 3 or Fig. 5. Correspondingly, during decoding, the prediction weights may for example be stored in an internal buffer of the Coefficient update block 212 of Fig. 4 or block 312 of Fig. 6, or in the channel buffer 202 of Fig. 4 or Fig. 6.
[0179] In the above discussion of modes 1, 2 and 3, reference has mainly been made to channels coded using LC prediction. It is however noted that initialization modes also may be used for initializing channels to be coded using both LP and LC prediction. It is noted that the initialization modes discussed above apply specifically to the sets of prediction weights used for the LC prediction. For the LP prediction, the set of prediction weights for the initial band n / o may typically be initialized with a set of zero- weight prediction weights, analogous to mode 1. For example, where the LP prediction weights are given by Eq. 5 or 5 ’, each prediction weight associated with the initial band0may be initialized according to a(nf,0, u) = 0 for each prediction weight index u.
[0180] In some implementations, the type of initialization mode may be varied on a frame-by-frame basis. For instance, mode 1 may as explained above be used when processing a frame for which information from a preceding frame is not available (e.g., since LC prediction was not used in the preceding frame), or where resetting of the predictor states is desired. Mode 2 may be used for any pair of consecutive frames where the prediction order for the cross-channel prediction is fixed and the prediction direction is altered from an upward prediction to a downward prediction direction, or vice versa. Mode 3 may be used where the prediction direction is fixed, or varied, and / or where the cross-channel prediction order varies between frames.
[0181] In some implementations, only one of modes 2 and 3 may be used in each frame of a sequence of frames where initialization based on information from a preceding frame is desired. Mode 1 may be used for every other frame of the sequence.
[0182] Regardless of whether the initialization modes are fixed or varied between frames, the encoding method may in some implementations further comprise encoding into the bitstreama flag indicating the type of initialization mode used (e.g., mode 1, mode 2 or mode 3), and / or a flag indicating the direction of the prediction to be used (i.e., an upward or downward prediction). The flag may be provided per channel of each frame The flag forms side information for, or of, its associated frame. The flag may facilitate prediction (e.g., LC and / or LP+LC prediction) during decoding, by allowing the decoder to readily identify the initialization mode and (where applicable) the prediction direction used during encoding by retrieving the flag from the bitstream and checking the value of the flag. Thus, the decoder may, for each frame / extract a flag associated with channel k and the frame / from the bitstream, and determine based on the flag what initialization mode and / or prediction direction to upon applying the predictive coding operation to a given channel k.
[0183] In implementations using only one of modes 2 and 3, the initialization mode may be signaled with a 1-bit flag, e.g., a first 1-bit value indicating initialization mode 1 and a second 1-bit value indicating initialization mode 2 or 3.
[0184] In implementations where the prediction direction may vary between frames, the prediction direction may be signaled using a 1-bit flag (e.g., a first 1-bit value indicating an upward prediction direction and a second 1-bit value indicating a downward prediction direction).
[0185] In implementations where the prediction direction is fixed, the prediction direction can be preconfigured for both the encoder and decoder and hence signaling of the prediction direction may be omitted.
[0186] In implementations where any one of modes 1, 2 and 3 may be used, the above- mentioned indications may be combined in a 2-bit flag with four possible states / values. For example: (first state) initialize prediction weights for channel k according to mode 1, use a downward prediction direction; (second state) initialize prediction weights for channel k according to mode 3, use a downward prediction direction; (third state) initialize prediction weights for channel k according to mode 2, use an upward prediction direction; (fourth state) initialize prediction weights for channel k according to mode 2, use a downward prediction direction.
[0187] It is noted that the either of the initialization modes discussed above in relation to the second frame f2 may be applied in a corresponding manner to the first frame f1, wherein accordingly, it is to be understood that the set of prediction weights for the initial band nf1,0 may be initialized based on prediction weights of a preceding frame to the first frame f1 (e.g., a frame f=f0=f1-1). That is, in the first frame f1, the set of prediction weights for the initial band nf1,0 and each given channel k may be initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band nf0,1 for the given channel k, in the preceding frame f0, wherein the final processed band nf0,1 in the preceding frame f0 is the same band as the initial processed band nf1,0 in the first frame f1;iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the preceding frame fo.
[0188] Systems and methods disclosed in the present application may be implemented as software, firmware, hardware or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; to the contrary, one physical component may have multiple functionalities, and one task may be carried out by several physical components in cooperation.
[0189] The computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware. Further, the present disclosure shall relate to any collection of computer hardware that individually or jointly execute instructions to perform any one or more of the concepts discussed herein.
[0190] Certain or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions that when executed by one or more of the processors carry out at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken are included. Thus, one example is a typical processing system (i.e. a computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system further may include a memory subsystem including a hard drive, SSD, RAM and / or ROM. A bus subsystem may be included for communicating between the components. The software may reside in the memory subsystem and / or within the processor during execution thereof by the computer system.
[0191] The one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s). Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0192] The software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitorymedia). As is well known to a person skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to the skilled person that communication media (transitory) typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0193] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the disclosure discussions utilizing terms such as “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer hardware or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.
[0194] It should be appreciated that in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of this invention. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0195] Furthermore, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means of carrying out the function. Thus, a processor with instructions forcarrying out such a method or element of a method forms a means for carrying out the method or element of a method. Note that when the method includes several elements, e.g., several steps, no ordering of such elements is implied, unless specifically stated. Furthermore, an element described herein of an apparatus embodiment is an example of a means for carrying out the function performed by the element for the purpose of carrying out the embodiments of the invention. In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0196] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0197] EEE 1. A method for encoding an input signal, the method comprising:obtaining a transform domain representation of a channel of a frame of the input signal, the transform domain representation including a set of input transform domain samples, wherein each input transform domain sample represents a respective band of the transform domain representation;applying a predictive coding operation to the channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, wherein the prediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual; andforming an encoded bitstream comprising encoded representations of each prediction residual.
[0198] EEE 2. A method for decoding an encoded bit stream, the method comprising: receiving the encoded bit stream, the encoded bit stream comprising encoded representations of a set of prediction residuals of a transform domain representation of a channel of a frame of an input signal, wherein each prediction residual represents a respective band of the transform domain representation;decoding the encoded bit stream to obtain the set of prediction residuals; andapplying a predictive coding operation to the channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, and wherein the reconstructed sample is determined based on the predicted sample and the prediction residual.
[0199] EEE 3. The method according to any one of the preceding EEEs, wherein the linear prediction of the predicted sample for each band comprises forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights, wherein an updated set of prediction weights is determined for each band.
[0200] EEE 4. The method according to EEE 3, wherein updating the set of prediction weights for a respective band comprises adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, by a respective step size.
[0201] EEE 5. The method according to EEE 4, wherein the respective step sizes for each respective band are determined based on the prediction residual of the previously processed band, and one or more of the reconstructed samples of the respective set of reconstructed samples.
[0202] EEE 6. The method according to EEE 5, wherein the respective step sizes are determined using a least mean squares method.
[0203] EEE 7. The method according to any one of the preceding EEEs, wherein the sequential processing of the bands proceeds in order from low to high frequency bands, or from high to low frequency bands.
[0204] EEE 8. The method according to any one EEEs 1-6, wherein the sequential processing of the bands proceeds in order from high to low frequency bands.
[0205] EEE 9. The method according to any one of the preceding EEEs, wherein each respective set of reconstructed samples comprises a set of reconstructed samples determined for previously processed bands for the channel.
[0206] EEE 10. The method according to EEE 9, wherein determining a predicted sample for a channel k and a band n comprises determining a linear prediction of PLP(k, ri),orwhere a(n, u) are prediction weights, X'(k,n + 1 + u) or X'(k,n — 1 — u) is a reconstructed sample of the respective set of reconstructed samples, and U is the number of reconstructed samples of the respective set of reconstructed samples.
[0207] EEE 11. The method according to any one of the preceding EEEs, when dependent on EEE 5, wherein each prediction weight a(n, u) is updated according toa(n — 1, u) = a(n, u) + D(k, ri) g X'(k, n + 1 + u) ora(n, u) = a(n + 1, u) + D(k,n — 1) g X'(k,n — 1 — u) where g is a gain factor.
[0208] EEE 12. The method according to EEE 11, wherein the gain factor g is given bywhere r < 1 and ||X'(k, n)|| is the norm of the respective set of previously reconstructed samples, and wherein, optionally, - — - — - is approximated by a right bit-shift of||X (k,n)||ceil{log2||X'(k, n)||2} >> 1.
[0209] EEE 13. The method according to EEE 1, or EEEs 3-12, when dependent on EEE 1, wherein the channel is a channel of a set of channels of the frame of the input signal, and wherein the method further comprises, for each of the further channels of the set of channels: obtaining a respective transform domain representation of the further channel of the frame, the transform domain representation including a respective set of input transform domain samples for the further channel, wherein each input transform domain sample represents a respective band of the transform domain representation; andapplying a predictive coding operation to the further channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, wherein the prediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual; andwherein the encoded bitstream is formed to further comprise encoded representations of each prediction residual for each further channel.
[0210] EEE 14. The method according to EEE 2, or any one EEEs 3-12, when dependent on EEE 2, wherein the channel is a channel of a set of channels of the frame of the input signal, wherein the encoded bit stream comprises encoded representations of a respective set of prediction residuals of a transform domain representation of each of the channels of the frame, wherein each prediction residual of the respective set of prediction residuals represents a respective band of a transform domain representation, and wherein the method further comprises, for each of the further channels of the set of channels:decoding the encoded bit stream to obtain the respective set of prediction residuals for the further channel; andapplying a predictive coding operation to the further channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, and wherein the reconstructed sample is determined based on the predicted sample and the prediction residual.
[0211] EEE 15. The method according to any one of EEEs 13-14, wherein the predictive coding operations are applied sequentially to the set of channels.
[0212] EEE 16. The method according to any EEEs 13-15, wherein, for each respective channel of the set of channels, the linear prediction of each respective predicted sample comprises an in-channel linear prediction based on a respective first set of reconstructed samples determined for previously processed bands for the respective channel, and / or a cross-channel linear prediction based on a respective second set of reconstructed samples determined for previously processed bands of a respective set of the other channels of the set of channels, for a same band as the respective predicted sample.
[0213] EEE 17. The method according to EEE 16, wherein, for each respective channel of the set of channels, the linear prediction of each respective predicted sample comprises an in-channel linear prediction based on a respective first set of reconstructed samples determined for previously processed bands, and cross-channel linear prediction based on a respective second set of reconstructed samples determined for previously processed bands of a respective subset of the set of channels, for a same band as the respective predicted sample.
[0214] EEE 18. The method according to any one of EEEs 16-17, wherein determining a predicted sample for a channel k of the set of channels and a band n comprises determining a cross-channel linear predictiwhere b(n, v) are prediction weights, X'(k — 1 — v, ri) is the set of reconstructed samples determined for a number of other channels of the set of channels, for the same band n as the respective predicted sample, and V is the number of the other channels.
[0215] EEE 19. The method according to EEE 18, when dependent on EEE 5, wherein each prediction weight is updated according tob(n — 1, v) = b(n, v) + D(k,ri) g X' (k — 1 — v, ri)or:b(n,v) = b(n — 1, v) + D(k,n — 1) g X'(k — 1 — v,n — 1) where g is a gain factor,and wherein, optionally, the gain factor g is given bywhere r < 1 and ||X'(k, n)|| is the norm of the respective set of previously reconstructed samples, and wherein, optionally, in determining the gain factor g, the term - —1is ||X (k,n) || approximated by a right bit-shift of ceiZ{log2|| "( / c, n) ||2}>> 1.
[0216] EEE 20. The method according to any one of EEEs 18-19, when dependent on EEE 10, wherein a predicted sample P(k, n) for a channel k and a band n is given byP(k,n) = PLPk,n) + PLC(k,n)
[0217] EEE 21. The method according to any one of EEEs 13-20, wherein, in each respective predictive coding operation, the respective set of the other channels of the set of channels is formed by a predetermined number of channels of the set of channels most recently subjected to a coding operation.
[0218] EEE 22. The method according to any one of EEEs 13-20, further comprising, determining an energy for each channel of the frame, wherein, in each respective predictive coding operation, the respective set of the other channels of the set of channels is formed by a predetermined number of the channels having the highest energy.
[0219] EEE 23. The method according to EEE 1, or any one of the EEEs dependent on EEE 1, further comprising:determining a prediction gain based on the set of input transform domain samples and the prediction residuals determined in the predictive coding operation for the channel, or, when dependent on EEE 13, a given channel of the set of channels; andcomparing the prediction gain to a prediction gain threshold to determine whether the encoded bit stream is to be formed to comprise encoded representations of the prediction residuals (or encoded representations of the input transform domain samples of the channel); wherein the encoded representations of the prediction residuals for the channel, or, when dependent on EEE 13, for each of the set of channels, are included in the encoded bitstream responsive to the prediction gain exceeding the prediction gain threshold.
[0220] EEE 24. The method according to EEE 23, wherein, responsive to the prediction gain exceeding the prediction gain threshold, forming the bit stream further comprises including in the encoded bitstream an indicator indicating that the channel, or each of the set of channels, has been encoded using linear predictive coding.
[0221] EEE 25. The method according to any one of EEEs 23-24, when dependent on EEE 16 or 17, further comprising determining, based on the prediction gain, whether to encode each channel of the set of channels using in-channel linear prediction, cross-channel linear prediction, or both.
[0222] EEE 26. The method according to EEE 2, or any one of the preceding EEEs dependent on EEE 2, wherein the encoded bit stream further comprises an indicator indicating whether the channel, or, when dependent on EEE 14, each channel of the set of channels, has been encoded using linear predictive coding, wherein the predictive coding operation is applied to the channel, or to each of the set of channels, responsive to the indicator indicating that the channel, or the given channel of the set of channels, has been encoded using linear predictive coding.
[0223] EEE 27. The method according to any one of EEEs 24-26, when dependent on EEE 16 or 17, wherein the indicator further indicates whether each channel has been encoded using in-channel linear prediction, cross-channel linear prediction, or both.
[0224] EEE 28. The method according to EEE 13, or any one of the preceding EEEs dependent on EEE 13, further comprising, for each channel of the set of channels:determining a respective prediction gain based on the set of input transform domain samples and the prediction residuals determined in the predictive coding operation for the channel; andcomparing the respective prediction gain to a prediction gain threshold to determine whether the encoded bit stream is to be formed to comprise encoded representations of the prediction residuals;wherein the encoded representations of the prediction residuals for the channel are included in the encoded bitstream responsive to the respective prediction gain exceeding the prediction gain threshold.
[0225] EEE 29. The method according to EEE 28, wherein, responsive to the respective prediction gain exceeding the prediction gain threshold, forming the bit stream further comprises including in the encoded bitstream an indicator indicating that the respective channel has been encoded using linear predictive coding.
[0226] EEE 30. The method according to any one of EEEs 28-29, when dependent on EEE 16 or 17, further comprising determining, based on the respective prediction gain, whether to encode the respective channel using in-channel linear prediction, cross-channel linear prediction, or both.
[0227] EEE 31. The method according to EEE 14, or any one of the preceding EEEs dependent on EEE 14, wherein the encoded bit stream further comprises, for each respective channel of the set of channels, a respective indicator indicating whether the respective channel has been encoded using in-channel linear prediction, wherein the predictive coding operation is applied to the respective channel responsive to the respective indicator indicating that the respective channel has been encoded using linear predictive coding.
[0228] EEE 32. The method according to EEE 29-31, when dependent on EEE 16 or 17, wherein the respective indicator further indicates whether the respective channel has been encoded using in-channel linear prediction, cross-channel linear prediction, or both.
[0229] EEE 33. The method according to EEE 1, or any one of the preceding EEEs dependent on EEE 1, wherein each input transform domain sample of the set of input transform domain samples represents a respective band of a first subset of bands of the transform domain representation, andwherein applying the predictive coding operation to the channel of the frame, comprises sequentially processing the first subset of bands, to determine, for each of the bands of the first subset of bands, the predicted sample, the prediction residual, and the reconstructed sample.
[0230] EEE 34. The method according to EEE 33, wherein the set of input transform domain samples is a first subset of input transform domain samples of the transform domain representation of the channel, and the transform domain representation further comprises a second subset of input transform domain samples for a second subset of bands of the transform domain representation,wherein applying the predictive coding operation to the channel of the frame further comprises sequentially processing the second subset of bands, to determine, for each of thebands of the second subset of bands, a predicted sample, a prediction residual, and a reconstructed sample, andwherein the encoded bitstream is formed to further comprise encoded representations of each prediction residual determined for the second subset of bands.
[0231] EEE 35. The method according to EEE 34, wherein for each band of the second subset of bands:the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands,the prediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample, andthe reconstructed sample is determined based on the predicted sample and the prediction residual.
[0232] EEE 36. The method according to any one of EEEs 33-35, wherein the set of input transform domain samples is a first subset of input transform domain samples of the transform domain representation of the channel, and the transform domain representation of the channel further comprises a third subset of input transform domain samples for a third subset of bands of the transform domain representation, andwherein the encoded bitstream is formed to further comprise encoded representations of the third subset of input transform domain samples.
[0233] EEE 37. The method according to EEE 36, wherein the encoded representation of each input transform domain samples of the third subset of input transform domain samples comprises the input transform domain sample or a quantized representation of the input transform domain sample.
[0234] EEE 38. The method according to EEE 2, or any one of the preceding EEEs dependent on EEE 2, wherein each prediction residual of the set of prediction residuals represents a respective band of a first subset of bands of the transform domain representation, andwherein applying the predictive coding operation to the channel of the frame, comprises sequentially processing the first subset of bands, to determine, for each of the bands of the first subset of bands, the predicted sample and the reconstructed sample.
[0235] EEE 39. The method according to EEE 38, wherein the set of prediction residuals is a first subset of prediction residuals of the transform domain representation of the channel, and the transform domain representation further comprises a second subset of prediction residuals for a second subset of bands of the transform domain representation,wherein applying the predictive coding operation to the channel of the frame further comprises sequentially processing the second subset of bands, to determine, for each of the bands of the second subset of bands, a predicted sample and a reconstructed sample.
[0236] EEE 40. The method according to EEE 39, wherein for each band of the second subset of bands:the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, andthe reconstructed sample is determined based on the predicted sample and the prediction residual.
[0237] EEE 41. The method according to any one of EEEs 39-40, wherein the transform domain representation of the channel further comprises a set of transform domain samples (e.g., non-prediction coded transform domain samples) for a third subset of bands of the transform domain representation, wherein each transform domain sample of the set of transform domain samples corresponds to an input transform domain sample of a transform domain representation of a channel of a frame of an input signal encoded in the bitstream.
[0238] EEE 42. The method according to EEE 41, wherein the encoded representation of each transform domain sample of the set of input transform domain samples comprises the input transform domain sample or a quantized representation of the input transform domain sample.
[0239] EEE 43. The method according to EEE 1, or any one of the EEEs dependent on EEE 1, wherein forming the bit stream comprises entropy encoding the prediction residuals.
[0240] EEE 44. The method according to EEE 1, or any one of the preceding EEEs dependent on EEE 1, wherein forming the bit stream further comprises quantizing each prediction residual.
[0241] EEE 45. The method according to EEE 1, or any one of the preceding EEEs dependent on EEE 1, wherein determining each prediction residual comprises determining a difference between the transform domain sample and the predicted sample, quantizing the difference and subsequently dequantizing the difference.
[0242] EEE 46. The method according to EEE 1, or any one of the preceding EEEs dependent on EEE 1, wherein the input signal is a waveform domain signal, wherein obtaining a transform domain representation of the channel, or each respective channel, of the frame comprises obtaining, for the channel or each respective channel, a sequence of waveform domain samples, and applying a frequency transform to each respective sequence of waveform domain samples.
[0243] EEE 47. The method according to EEE 46, wherein the input signal is a time domain signal and each waveform domain sample is a time domain sample.
[0244] EEE 48. The method according to any one of EEEs 46-47, wherein the frequency transform is a discrete Fourier transform, a discrete cosine transform (DCT), such as DCT Type II or Modified DCT (MDCT), or a discrete sine transform (DST).
[0245] EEE 49. A method for encoding an input signal comprising a set of channels, the method comprising, for each of a first frame, f=fi, and a second frame, f=f2, of the input signal:obtaining a transform domain representation of the frame / , the transform domain representation including, for each channel, k, of the set of channels, a set of transform domain samples, each transform domain sample / ( / c, n) associated with a respective band, n, of the transform domain representation;applying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples, a set of prediction residuals, and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, nf,0, to a final band, nf,1, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ;a prediction residual, Df(k, n), is determined based on the transform domain sample Xf (k, n) and the predicted sample Pf(k, n) determined for the given band, n, and represents a prediction error between the transform domain sample X (k, n) and the predicted sample Py (k, n); anda reconstructed sample, X' f (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) determined for the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each prediction weight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band np.o in the second frame / 2, the set of prediction weights is initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band n / 7,7 for the given channel k, in the first frame fi, wherein the final processed band n / 7,7 in the first frame / / is the same band as the initial processed band n / 2,oin the second frame f,iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f,and the method further comprising encoding representations of the prediction errors determined for each band n and channel k into a bitstream.
[0246] EEE 50. A method for decoding an encoded bitstream, the method comprising: decoding the encoded bitstream to obtain, for each of a first frame, f=fi, and a second frame, = 2, a set of prediction residuals for each channel k of a set of channels, each prediction residual, Df(k, n), associated with a respective band, n, of a transform domain representation;and for each of the first frame / / and the second frame f.applying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, n / o, to a final band, n 7, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ; anda reconstructed sample, X'f (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) associated with the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each prediction weight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band nf2,0 in the second frame f2, the set of prediction weights is initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band n / 7,7 for the given channel k, in the first frame fi, wherein the final processed band n / 7,7 in the first frame fi is the same band as the initial processed band n / z.oin the second frame 2;iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f1.
[0247] EEE 51. The method according to any one of EEEs 49-50, wherein determining the predicted sample, Pf(k, n) comprises determining a cross-channel linear prediction of PLC(, ri) according to:where bfk(n, v) is a prediction weight, X (k — 1 — v, n) is a reconstructed sample, and V is the number of previously reconstructed samples.
[0248] EEE 52. The method according to EEE 51, wherein:where rif^ < rifo each prediction weight is updated according tofor each band> n^Q, each prediction weight is updated according tobf,k (n+ 1, v) = bfk(n, v) + Df (k, ri) g Xf(k — 1 — v, ri) for each bandwhere g is a gain factor.
[0249] EEE 53. The method according to EEE 52, wherein the gain factor g is given bywhere r < 1 andis the norm of the respective set of previously reconstructed samples, and wherein, optionally, in determining the gain factor g, the term..1— n- is ||X / (fc,n)||2approximated by a right bit-shift of cei / {log2|| ^ (k, ri) || }>>1.
[0250] EEE 54. The method according to any one of EEEs 49-53, wherein, for the given channel k and the initial band nj2,o in the second frame f2, the set of prediction weights is initialized according to option iii.
[0251] EEE 55. The method according to EEE 54, when dependent on EEE 51, wherein each given prediction weightinitialized according towhere N is at least a subset of the bands from the initial band ng,o to the final band ng in the first frame / ; and Ntotis the number of bands of the at least a subset of bands.
[0252] EEE 56. The method according to any one of EEEs 54-55, whereinfor each of the first and second frames / ; and / 2.
[0253] EEE 57. The method according to any one of EEEs 49-53, wherein, for the given channel k and the initial band nf2,0 in the second frame f2, the set of prediction weights is initialized according to option ii.
[0254] EEE 58. The method according to EEE 57, wherein the first and second frames / ;, f2 are comprised in an alternating sequence of first frames and second frames,wherein, for each first frame n / o < n / i, and for each second frame n / o > n / , wherein for each second frame / the initial band n o is the same band as the final band n o of its preceding frame / - 1;wherein each frame, / of the first and second frames are subjected to a predictive coding operation as set out in EEE 49 or EEE 50; andwherein, for each given frame / the set of prediction weights for the initial band n.f, for each given channel k, is initialized with the set of prediction weights used for processing the final band n / ,; for the given channel k, in its preceding frame f-1.
[0255] EEE 59. The method according to any one of EEEs 49-58, wherein, for at least one channel k, the linear prediction of each predicted sample Pf(k, n), is determined as a sum of: the cross-channel linear prediction, PfLC(k> n), and an in-channel linear prediction, PfLP(k, n), based on a respective first set of previously reconstructed samples determined for previously processed bands for the channel k.
[0256] EEE 60. The method according to EEE 59, wherein determining the in-channel linear prediction for a channel k and a band n comprises determining:where the sequential processing of the bands proceeds in order from high to low frequency bands,or:where the sequential processing of the bands proceeds in order from low to high frequency bands,where a(n, u) are prediction weights, Xf(k,n + 1 + u) or Xf(k,n — 1 — u) is a reconstructed sample of the respective set of reconstructed samples, and U is the number of reconstructed samples of the respective set of reconstructed samples.
[0257] EEE 61. The method according to EEE 60, wherein the prediction weights a(n, u) are updated according to:a(n — 1, u) = a(n, u) + D f(k, n) g X'f(k, n + 1 + u), where the sequential processing of the bands proceeds in order from high to low frequency bands,or:a(n + 1, u) = a(n, u) + D f(k, n) g X'f(k, n — 1 — u) where the sequential processing of the bands proceeds in order from low to high frequency bands,where g is a gain factor.
[0258] EEE 62. The method according to EEE 61, wherein the gain factor g is given bywhere r < 1 and ||X'(k, n)|| is the norm of the respective set of previously reconstructed samples, wherein, optionally, in determining the gain factor g, the term — — r is2approximated by a right bit-shift of ceil{log2||X'f(k, n)||2}>>1.
[0259] EEE 63. The method according to any one of EEEs 49-62, when dependent on EEE 49, further comprising, for the second frame f2, encoding into the bitstream a flag indicating, for each given channel k, a type of initialization used for initializing the set of prediction weights for the initial band np,o, and / or whether the initial band np.o is a higher or lower band than the final band, np,i.
[0260] EEE 64. The method according to any one of EEEs 49-62, when dependent on EEE 50, further comprising, for the second frame f2, retrieving from the encoded bitstream a flag indicating, for each given channel k, a type of initialization used for initializing the set of prediction weights for the initial band np,0 during encoding, and / or whether the initial band np,0 is a higher or lower band than the final band, np,1.
[0261] EEE 65. The method according to any one of EEEs 49-64, wherein determining each prediction residual comprises determining a difference between the input transform domain sample and the predicted sample, quantizing the difference and subsequently dequantizing the difference.
[0262] EEE 66. The method according to EEE 65, wherein the encoded representation of each prediction residual is an entropy coded representation of the quantized difference between the input transform domain sample and the predicted sample.
[0263] EEE 67. A device comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of the preceding EEEs.
[0264] EEE 68. A computer program product comprising computer program code portions configured to perform the method according to one of EEEs 1-66 when executed on a computer processor.
Claims
CLAIMS1. A method for encoding an input signal, the method comprising:obtaining a transform domain representation of a channel of a frame of the input signal, the transform domain representation including a set of input transform domain samples, wherein each input transform domain sample represents a respective band of the transform domain representation;applying a predictive coding operation to the channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, the linear prediction comprising forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights,wherein the prediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual, andwherein an updated set of prediction weights is determined for each band, and wherein updating the set of prediction weights for a respective band comprises adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, by a respective step size; andforming an encoded bitstream comprising encoded representations of each prediction residual.
2. A method for decoding an encoded bit stream, the method comprising:receiving the encoded bit stream, the encoded bit stream comprising encoded representations of a set of prediction residuals of a transform domain representation of a channel of a frame, wherein each prediction residual represents a respective band of the transform domain representation;decoding the encoded bit stream to obtain the set of prediction residuals; and applying a predictive coding operation to the channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, the linear prediction comprising forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual,wherein an updated set of prediction weights is determined for each band, and wherein updating the set of prediction weights for a respective band comprises adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, by a respective step size.
3. The method according to any one of the preceding claims, wherein the respective step sizes for each respective band are determined based on the prediction residual of the previously processed band, and one or more of the reconstructed samples of the respective set of reconstructed samples.
4. The method according to any one of the preceding claims, wherein the respective step sizes are determined using a least mean squares method.
5. The method according to any one of the preceding claims, wherein the sequential processing of the bands proceeds in order from low to high frequency bands, or from high to low frequency bands.
6. The method according to any one of the preceding claims, wherein each respective set of reconstructed samples comprises a set of reconstructed samples determined for previously processed bands for the channel.
7. The method according to claim 6, wherein determining a predicted sample for a channel k and a band n comprises determining a linear prediction of PLP(k, ri),where the sequential processing of the bands proceeds in order from high to low frequency bands, orwhere the sequential processing of the bands proceeds in order from low to high frequency bands,where a(n, u) are prediction weights, X'(k, n + 1 + u) or X'(k, n — 1 — u) is a reconstructed sample of the respective set of reconstructed samples, and U is the number of reconstructed samples of the respective set of reconstructed samples.
8. The method according to claim 7, wherein the prediction weights a(n, u) are updated according to:a(n — 1, u) = a(n, u) + D(k, n) g X'(k, n + 1 + u), where the sequential processing of the bands proceeds in order from high to low frequency bands,or:a(n + 1, u) = a(n, u) + D(k, n) g X'(k, n — 1 — u) where the sequential processing of the bands proceeds in order from low to high frequency bands,where g is a gain factor.
9. The method according to claim 8, wherein the gain factor g is given bywhere r < 1 and ||X'(k, n)|| is the norm of the respective set of previously reconstructed samples, wherein, optionally, in determining the gain factor g, the term 1 / ||X'(k,n)|| is approximated by a right bit-shift of ceil{log2||X'(k, n)||2}>>1.
10. The method according to claim 1, or any one of claims 3-9, when dependent on claim 1, wherein the channel is a channel of a set of channels of the frame of the input signal, and wherein the method further comprises, for each of the further channels of the set of channels: obtaining a respective transform domain representation of the further channel of the frame, the transform domain representation including a respective set of input transform domain samples for the further channel, wherein each input transform domain sample represents a respective band of the transform domain representation; andapplying a predictive coding operation to the further channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, the linear prediction comprising forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights,wherein the prediction residual is determined based on the input transform domain sample and the predicted sample and represents a prediction error between the input transform domain sample and the predicted sample, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual, andwherein an updated set of prediction weights is determined for each band, and wherein updating the set of prediction weights for a respective band comprises adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, for the further channel, by a respective step size; andwherein the encoded bitstream is formed to further comprise encoded representations of each prediction residual for each further channel.
11. The method according to claim 2, or any one of claims 3-9, when dependent on claim 2, wherein the channel is a channel of a set of channels of the frame, wherein the encoded bit stream comprises encoded representations of a respective set of prediction residuals of a transform domain representation of each of the channels of the frame, wherein each prediction residual of the respective set of prediction residuals represents a respective band of the transform domain representation, and wherein the method further comprises, for each of the further channels of the set of channels:decoding the encoded bit stream to obtain the respective set of prediction residuals for the further channel; andapplying a predictive coding operation to the further channel of the frame, comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample and a reconstructed sample,wherein the predicted sample is determined using linear prediction based on a respective set of reconstructed samples determined for previously processed bands, thelinear prediction comprising forming a weighted combination of the respective set of reconstructed samples using a set of prediction weights, andwherein the reconstructed sample is determined based on the predicted sample and the prediction residual,wherein an updated set of prediction weights is determined for each band, and wherein updating the set of prediction weights for a respective band comprises adjusting each prediction weight of the set of prediction weights used for a previously processed band neighboring to the respective band, for the further channel, by a respective step size.
12. The method according to any one of claims 10-11, wherein the predictive coding operations are applied sequentially to the set of channels.
13. The method according to any one of claims 10-12, wherein, for each respective channel of the set of channels, the linear prediction of each respective predicted sample comprises an in-channel linear prediction based on a respective first set of reconstructed samples determined for previously processed bands for the respective channel, and / or a cross-channel linear prediction based on a respective second set of reconstructed samples determined for previously processed bands of a respective set of the other channels of the set of channels, for a same band as the respective predicted sample.
14. The method according to claim 13, wherein, for at least one of the set of channels, the linear prediction of each respective predicted sample comprises a cross-channel linear prediction based on a respective second set of reconstructed samples determined for previously processed bands of a respective set of the other channels of the set of channels, for a same band as the respective predicted sample.
15. The method according to claim 14, wherein determining a cross-channel linear prediction P LC(k, n) of a predicted sample for a channel k of the set of channels and a band n comprises determiningwhere b(n, v) are prediction weights, X'(k — 1 — v, n) is the set of reconstructed samples determined for a number of other channels of the set of channels, for the same band n as the respective predicted sample, and V is the number of the other channels.
16. The method according to claim 15, wherein the prediction weights b(n, v) are updated according to:b(n — 1, v) = b(n, v) + D (k,ri) g X' (k — 1 — v, n)where the sequential processing of the bands proceeds in order from high to low frequency bands,or:b(n,v) = b(n — 1, v) + D(k,n — 1) g X'(k — 1 — v,n — 1) where the sequential processing of the bands proceeds in order from high to low frequency bands,where g is a gain factor.
17. The method according to 16, wherein the gain factor g is given bywhere r < 1 and 11 X' (k, ri) || is the norm of the respective set of previously reconstructed samples, and wherein, optionally, in determining the gain factor g, the term - — - — - is ||X (k,n) || approximated by a right bit-shift of ceil{log2||X'(k, n)||2}>>1.
18. The method according to any one of claims 15-17, when dependent on claim 13, wherein a predicted sample P(k, n) for a channel k and a band n is given by:P(k, n) = PLP(k, n) + PLC(k, n).
19. The method according to claim 1, or any one of the preceding claims dependent on claim 1, further comprising:determining a prediction gain based on the set of input transform domain samples and the prediction residuals determined in the predictive coding operation for the channel, or, when dependent on claim 10, a given channel of the set of channels; andcomparing the prediction gain to a prediction gain threshold to determine whether the encoded bit stream is to be formed to comprise encoded representations of the prediction residuals;wherein the encoded representations of the prediction residuals for the channel, or, when dependent on claim 10, for each of the set of channels, are included in the encoded bitstream responsive to the prediction gain exceeding the prediction gain threshold, andwherein, optionally, responsive to the prediction gain exceeding the prediction gain threshold, forming the bit stream further comprises including in the encoded bitstream anindicator indicating that the channel, or each of the set of channels, has been encoded using linear predictive coding.
20. The method according to claim 2, or any one of the preceding claims dependent on claim 2, wherein the encoded bit stream further comprises an indicator indicating whether the channel, or, when dependent on claim 11, each channel of the set of channels, has been encoded using linear predictive coding, wherein the predictive coding operation is applied to the channel, or to each of the set of channels, responsive to the indicator indicating that the channel, or each channel of the set of channels, has been encoded using linear predictive coding.
21. The method according to claim 1, wherein the input signal comprises a first frame, f=fi, and a second frame, f=f2, each comprising a set of channels, wherein the frame is the second frame f2, and wherein the method comprises, for each of the first frame f1 and the second frame f2: obtaining a transform domain representation of the frame the transform domain representation including, for each channel, k, of the set of channels, a set of transform domain samples, each transform domain sample / ( / c, n) associated with a respective band, n, of the transform domain representation;applying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples, a set of prediction residuals, and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, nf,0, to a final band, nf,1, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ;a prediction residual, Df(k, n), is determined based on the transform domain sample Xf (k, n) and the predicted sample Pf(k, n) determined for the given band, n, and represents a prediction error between the transform domain sample Xf(k, n) and the predicted sample Pf(k, n); anda reconstructed sample, X' f (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) determined for the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each prediction weight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band nf2,0 in the second frame f2, the set of prediction weights is initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band nf1,1 for the given channel k, in the first frame f1, wherein the final processed band nf1,1 in the first frame f1 is the same band as the initial processed band nf2,0 in the second frame f2;iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f1.
22. The method according to claim 2, wherein the encoded bit stream comprises encoded representations of a set of prediction residuals of each of a first frame, f=f1, and a second frame, f=f2, each comprising a set of channels, wherein the frame is the second frame / 2, and wherein the method comprises, for each of the first frame f1 and the second frame f2:decoding the encoded bit stream to obtain the set of prediction residuals for each channel, k, of the set of channels, each prediction residual, Df(k, n), associated with a respective band, n, of a transform domain representation; andapplying a predictive coding operation to each channel k of the set of channels, to determine a set of predicted samples and a set of reconstructed samples for each channel, wherein for a given channel, k, the predictive coding operation comprises sequentially processing the bands, in a direction from an initial band, nf,0, to a final band, nf,1, wherein, for the given channel k and a given band, n:a predicted sample, Pf(k, n), is determined using a cross-channel linear prediction based on a set of prediction weights associated with the given channel k, and a respective set of previously reconstructed samples determined for the given band n, for a set of previously coded channels of the frame / ; anda reconstructed sample, X' (k, n), is determined based on the predicted sample Pf(k, n) and the prediction residual Df(k, n) associated with the given band n; andwherein, for the given channel k and each given band, n, except the initial band nf,0, an updated set of prediction weights is determined for the given band n, wherein each predictionweight of the updated set of prediction weights is determined based on a corresponding prediction weight of a set of prediction weights used for a previously processed band, m, neighboring to the given band n, and a respective step size; andwherein, for the given channel k and the initial band nf2,0 in the second frame f2, the set of prediction weights is initialized according to one of the following options:i. with a set of zero- valued prediction weights;ii. with the set of prediction weights used for processing the final band ng,] for the given channel k, in the first frame;, wherein the final processed band ng in the first frame / ; is the same band as the initial processed band n / 2,oin the second frame / 2;iii. each given prediction weight of the set of prediction weights is initialized with an average prediction weight value computed as an average of at least a subset of its corresponding prediction weights determined for the channel k in the first frame f1.
23. The method according to any one of claims 21-22, wherein determining the predicted sample, Pf(k, n) comprises determining a cross-channel linear prediction of PLC(, ri) according to:where bfk(n, v) is a prediction weight, Xf(k — 1 — v, n) is a reconstructed sample, and V is the number of previously reconstructed samples.
24. The method according to claim 23, wherein:where rif^ < ngQ, each prediction weight is updated according tobyk(n — 1, v) = byk(n, v) + y (k, ri) g Xf(k — 1 — v, ri) for each band> n^Q, each prediction weight is updated according tofor each band n = ngQngQ+ 1, ny0+ 2,..., ny15where g is a gain factor.
25. The method according to claim 24, wherein the gain factor g is given bywhere r < 1 andis the norm of the respective set of previously reconstructed samples, and wherein, optionally, in determining the gain factor g, the term..,1— n- is ||X / (fc,n)||2approximated by a right bit-shift of cei / {log2|| ^ (k, ri) || }>>1.
26. The method according to any one of claims 21-25, wherein, for the given channel k and the initial band nj2,o in the second frame f2, the set of prediction weights is initialized according to option iii.
27. The method according to claim 26, when dependent on claim 23, wherein each given prediction weight bf2,knf2,o>v) isinitialized according to1bf2,k (nf2,0>V) y bfi,k n, v').where N is at least a subset of the bands from the initial band n;,oto the final bandin the first frame; and Ntotis the number of bands of the at least a subset of bands.
28. The method according to any one of claims 26-27, wherein rif^for each of the first and second frames / ; and / 2.
29. The method according to any one of claims 21-25, wherein, for the given channel k and the initial band np.o in the second frame / 2, the set of prediction weights is initialized according to option ii.
30. The method according to claim 29, wherein the first and second frames / ;, / ? are comprised in an alternating sequence of first frames and second frames,wherein, for each first frame n / o < n / i, and for each second frame n / o > n / , wherein for each second frame / the initial band n o is the same band as the final band n of its preceding frame f-1;wherein each frame, / of the first and second frames are subjected to a predictive coding operation as set out in claim 21 or 22; andwherein, for each given frame / the set of prediction weights for the initial band n.f, for each given channel k, is initialized with the set of prediction weights used for processing the final band n / 7,7 for the given channel k, in its preceding frame f-1.
31. The method according to any one of claims 21-30, wherein, for at least one channel k, the linear prediction of each predicted sample Pf(k, n), is determined as a sum of: the cross-channel linear prediction, PfLC(k,n) and an in-channel linear prediction, PfLP(k,n) based on a respective first set of previously reconstructed samples determined for previously processed bands for the channel k.
32. The method according to any one of the preceding claims, wherein determining each prediction residual comprises determining a difference between the input transform domain sample and the predicted sample, quantizing the difference and subsequently dequantizing the difference.
33. The method according to claim 32, wherein the encoded representation of each prediction residual is an entropy coded representation of the quantized difference between the input transform domain sample and the predicted sample.
34. A device comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of the preceding claims.
35. A computer program product comprising computer program code portions configured to perform the method according to one of claims 1-33 when executed on a computer processor.
Citation Information
Patent Citations
Prediction of spectral coefficients in waveform coding and decoding
US20070016415A1
US202463713197P
US202463713209P