Mutually exclusive use of temporal noise shaping mode or noise synthesis mode
The decoder and encoder system efficiently applies TNS and NS modes with a single set of filter parameters, addressing the balance between signal quality, computational costs, and bitrate demands in perceptual coding by mutually exclusive mode switching.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2025-10-23
- Publication Date
- 2026-04-30
AI Technical Summary
Existing general-purpose perceptual coding of digital waveform signals faces challenges in achieving a balance between signal quality, computational costs, and bitrate demands due to the significant requirements for transmitting temporal noise shaping and noise synthesis parameters.
A decoder and encoder system that mutually exclusively applies temporal noise shaping (TNS) and noise synthesis (NS) modes, using a single set of filter parameters and predictive or non-predictive mode switching to reduce signaling effort and computational costs, while maintaining signal quality.
This approach reduces bitrate and computational load by allowing TNS and NS functionalities to be applied in a mutually exclusive manner, optimizing the trade-off between quality and efficiency in encoding and decoding digital waveform signals.
Smart Images

Figure EP2025080596_30042026_PF_FP_ABST
Abstract
Description
[0001] MUTUALLY EXCLUSIVE USE OF TEMPORAL NOISE SHAPING MODE
[0002] OR NOISE SYNTHESIS MODE
[0003] Description
[0004] Technical Field
[0005] Embodiments comprise decoders, encoders, methods for decoding, methods for encoding, computer programs and data streams for an efficient signaling of noise synthesis and shaping parameters for general-purpose perceptual waveform coding.
[0006] Background of the Invention
[0007] General-purpose perceptual coding of digital waveform signals, such as biophysical (e.g., medical), geophysical (e.g., seismic), or acoustic (e.g., audio) signals comprises, or for example even requires, three types of basic functionality:
[0008] 1. spectral shaping of the quantization distortion, e.g. most commonly or optionally by filtering of the (re)quantized decoded signal coefficients, optionally, with a linear predictive coding (LPC)-like, for example spectrally "non-flat" filter; this functional component is often referred to as "spectral noise shaping", abbreviated SNS here,
[0009] 2. temporal shaping of the quantization distortion, e.g. either via time-domain gain control with scalar multiplication of the decoded time-domain signal with transmission of the scalar values, or e.g. via temporal noise shaping (TNS) of transformed versions of the decoded signal coefficients; see[1],
[0010] 3. synthesis of noise-like components (for example, most commonly background or recording noise) missing in the decoded signals, e.g., due to coarse quantization of said transform - domain signal coefficients; see [2],
[0011] The SNS functional component is well studied and may be based on or may for example simply require time-domain synthesis filtering of the decoded (i.e., frame-wise reconstructed) waveform signal, for example preferably with time-varying filters.
[0012] However, a use of temporal shaping of the quantization distortion and a use of noise synthesis may cause a significant demand in bitrate for transmitting the respective shaping and synthesis parameters. Therefore, it is desired to obtain a coding concept (e.g. for encoding and respectively decoding of a digital waveform signal) which achieves a better compromise between a quality of an encoded and subsequently decoded signal, computational costs, and bitrate demands.
[0013] This is achieved by the subject matter of the independent claims of the present application.
[0014] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application.
[0015] of the invention
[0016] An embodiment according to the invention comprises a decoder for decoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) from a data stream using block-based transform decoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), wherein the decoder is configured to decode a spectrum (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum) for a current (e.g. a currently considered block, e.g. a currently encoded block) block of the DWS from the data stream, decode a mode signal (e.g. percept_mode) from the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the decoder is configured to subject the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) (e.g. to decode the current block by subjection the spectrum to TNS), in case of the selected mode being the TNS mode; and wherein the decoder is configured to subject the spectrum to NS (e.g. to decode the current block by subjection the spectrum to NS), in case of the selected mode being the NS mode. It was recognized that temporal noise shaping and noise synthesis may, at least for some applications, be used or applied preferably in a mutually exclusive fashion. As an example, it was recognized that, for frames or blocks, which may be efficiently or effectively coded using temporal noise shaping, an application of noise synthesis may have only little advantages or impact on the quality of the reconstructed signal (e.g. after encoding and decoding), and vice versa that for frames or blocks, which may be efficiently or effectively coded using noise synthesis, an application of temporal noise shaping may have only little advantages or impact for the quality of the reconstructed signal (e.g. after encoding and decoding).
[0017] Hence, it was recognized that the TNS and NS functionality may be applied in a mutually exclusive manner, hence reducing signaling effort and computational costs.
[0018] According to an embodiment of the invention, the decoder is configured to decode filter parameters from the data stream for the current block (e.g. irrespective of the selected mode being the NS mode or the TNS mode, e.g. in a mode agnostic manner) (e.g. wherein the decoder is configured to decode a noise level information form the data stream, in case of the selected mode being the NS mode); wherein the decoder is configured to use the filter parameters for the TNS filtering, in case of the selected mode being the TNS mode; andwherein the decoder is configured to fill zero-quantized portions of the spectrum using noise, and shape the noise using the filter parameters (e.g. wherein the decoder is configured to use the noise level information form the data stream, in order to perform the noise synthesis), in case of the selected mode being the NS mode.
[0019] It was recognized that it may be sufficient to provide decoders having both temporal noise shaping and noise synthesis functionalities with only one set of filter parameters (e.g. instead of two distinct sets of filter parameters) via the bitstream. As the functionalities may be applied in a mutually exclusive manner, the load on the bitstream can be reduced by only having to provide only one set of filter parameters for use in either the temporal noise shaping and noise synthesis functionality.
[0020] According to an embodiment of the invention, the decoder is configured to decode the filter parameters by predicting the filter parameters based on filter parameters signaled in the data stream for a previous block ((e.g. a previously decoded block of the DWS, e.g. a previously transmitted block of the DWS) to obtain predicted filter parameters, decoding a prediction residual ((e.g. based on the predicted filter parameters and the filter parameters for the current block), and correcting the predicted filter parameters using the prediction residual. Hence, the filter parameters may be obtained in a particularly efficient manner based on a prediction processing, e.g. a decoder-sided prediction processing, e.g. a prediction processing which is an inverse prediction processing compared to an encoder-sided prediction processing, e.g. by applying an inverse filter functionality compared to the encoder-side, e.g. by reconstructing redundant signal portions, which were reduced by the encoder-sided prediction processing.
[0021] According to an embodiment of the invention, the decoder is configured to decode the filter parameters by predicting the filter parameters based on filter parameters signaled in the data stream for the previous block (e.g. previously decoded block of the DWS, e.g. e. previously transmitted block of the DWS) to obtain the predicted filter parameters, decoding the prediction residual and correcting the predicted filter parameters using the prediction residual, irrespective of whether the filter parameters for the current block and the filter parameters signaled in the data stream for the previous block are associated with a same selected mode (e.g. if the previous block is coded using the TNS mode and if the previous block is coded using the NS mode (e.g. irrespective of whether TNS mode or NS mode is selected) (e.g. irrespective of whether the current block and the previous block are associated with a same selected mode).
[0022] Hence, it was recognized that the filter parameter prediction may be performed irrespective of whether TNS parameters are predicted from previous NS parameters or vice versa. It was recognized that by predicting filter parameters even for different functionalities, a load on the bitstream can be reduced, e.g. compared to a provision of two separate sets of filter coefficients.
[0023] According to an embodiment of the invention, the decoder is configured to decode a parameter mode information (e.g. delta_time_flag) from the data stream, which indicates a selected parameter mode out of a predictive mode (e.g. delta_time_flag having value 1) and a non-predictive mode (e.g. delta_time_flag having value 0); decode the filter parameters by predicting the filter parameters based on filter parameters signaled in the data stream for the previous block of the DWS to obtain the predicted filter parameters, decoding the prediction residual and correcting the predicted filter parameters using the prediction residual, in case of the selected parameter mode being the predictive mode; decode the filter parameters (e.g. in a non-differential manner, e.g. without predictive model or computing residuals) from the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS, in case of the selected parameter mode being the non-predictive mode. Hence, given such a switching functionality, it can be selectively decided whether a predictionbased coding or an independent coding is advantageous, e.g. with respect to a computational load and / or a number of bits to be encoded.
[0024] According to an embodiment of the invention, the the parameter mode information (e.g. delta_time_flag) is a block-specific and / or channel specific information (e.g. wherein the DWS comprises a plurality of channels); and
[0025] wherein the decoder is configured to selectively decode the filter parameters using the predictive mode or the non-predictive mode on a per-block basis and / or on a per-channel basis.
[0026] Hence, a granularity of the switching may be adapted, e.g. allowing to reduce a signaling frequency of the parameter mode information.
[0027] According to an embodiment of the invention, the decoder is configured to decode the filter parameters as reflection coefficients and / or as lattice coefficients.
[0028] It was recognized that such a form of filter parameters may allow a use as both TNS filter coefficients as well as noise shaping filter coefficients.
[0029] According to an embodiment of the invention, the decoder is configured to perform the prediction of the filter parameters and / or the correction of the predicted filter parameters in a spectral domain (in other words, as an example, even prediction and / or correction may be done in this representation domain, e.g. corresponding to the reflection coefficients and / or as lattice coefficients).
[0030] It was recognized that such a processing in spectral domain is efficient.
[0031] According to an embodiment of the invention, the decoder is configured to subject the spectrum to TNS filtering, in case of the selected mode being the TNS mode and subject the spectrum to NS, in case of the selected mode being the NS mode, using a same filter structure (e.g. a prediction filtering) (e.g. a same filtering structure but with different filter orders and / or different filter parameters). Using a same filter structure may allow reducing a complexity of the decoder and simplify a reusability (e.g. prediction of TNS parameters from NS parameters and vice versa) of the filter parameters.
[0032] According to an embodiment of the invention, the decoder is configured to decode noise synthesis parameters (e.g. comprising noise level and spectral shape parameters; e.g. comprising an information about a noise level energy and a spectral tilt, e.g. noise synthesis parameters which are not filter parameters) from the data stream for the current block; fill zero-quantized portions of the spectrum using noise, and shape the noise using the noise synthesis parameters, in case of the selected mode being the NS mode.
[0033] According to an embodiment of the invention, the decoder is is configured to selectively switch (e.g. on a block-wise basis, e.g. on a frame-wise basis, e.g. on a channel-wise basis) between decoding filter parameters, and using the filter parameters for the TNS filtering, and performing time-domain gain control in order to perform TNS (e.g. with scalar multiplication of the decoded time-domain signal with transmission of the scalar values) (e.g. in order to subject the spectrum to TNS filtering), in case of the selected mode being the TNS mode.
[0034] It was recognized that for some cases a switching to time-domain gain control may be advantageous, hence, such a switching functionality may increase an efficiency of the decoder.
[0035] According to an embodiment of the invention, the decoder is configured to decode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode (e.g. TNS mode, e.g. NS mode, e.g. when having a value other than zero, e.g. a value of 2), which is to be considered (e.g. applied) for every block of the DWS.
[0036] According to an embodiment of the invention, the decoder is configured to decode a channelspecific high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode (e.g. the TNS mode, e.g. the NS mode, e.g. when having a value other than zero, e.g. a value of 2), which is to be considered (e.g. applied) for every block of a channel (or at least of a portion of this channel) of the DWS (e.g. wherein the DWS comprises a plurality of channels).
[0037] According to an embodiment of the invention, the decoder is configured to decode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is any out of the TNS mode or the NS mode (e.g. indicating whether the selected mode is a perceptual mode, e.g. comprising spectral shaping on the basis of perceptual characteristics); and wherein the decoder is configured to skip decoding the mode signal, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode (e.g. when percept_mode 2).
[0038] It was recognized that a high level syntax element provides for efficient means for the respective signaling.
[0039] According to an embodiment of the invention, the mode signal comprises a m-ary syntax element (e.g. with m>=2), wherein a first value indicates the TNS mode and wherein a second value indicates the NS mode.
[0040] According to an embodiment of the invention, the decoder is configured to decode filter parameters from the data stream for the current block (e.g. irrespective of the selected mode being the NS mode or the TNS mode, e.g. in a mode agnostic manner); wherein the decoder is configured to decode a noise level information (e.g. a noise level energy), in case of the selected mode being the NS mode, wherein the decoder is configured to use the filter parameters for the TNS filtering, in case of the selected mode being the TNS mode; and wherein the decoder is configured to fill zero-quantized portions of the spectrum using noise and shape the noise using the filter parameters and using the noise level information, in case of the selected mode being the NS mode.
[0041] An embodiment according to the invention comprises an encoder for encoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) into a data stream using block-based transform encoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), wherein the encoder is configured to encode a current block (e.g. a currently considered block, e.g. a currently encoded block) of the DWS into the data stream by deriving a spectrum of the current block (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum), encode a mode signal (e.g. percept_mode) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the encoder is configured to, in encoding the current block, encode the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) (e.g. decoding the current block by subjection the spectrum to TNS filtering), in case of the selected mode being the TNS mode; and subjecting the spectrum to NS (e.g. decoding the current block by subjection the spectrum to NS), in case of the selected mode being the NS mode.
[0042] The encoder according to embodiments may be based on the same principles as the corresponding decoder. Hence, the previous advantages, features, functionalities and details discussed above for decoder-aspects may be understood in a corresponding manner or implemented in a corresponding manner, individually or taken in combination, in a respective encoder.
[0043] According to an embodiment of the invention, the encoder is configured to encode filter parameters into the data stream for the current block (e.g. in a manner so that the decoding of the filter parameters may be done irrespective of the selected mode being the NS mode or the TNS mode, e.g. in a mode agnostic manner; in other words, the “irrespective” may relate to the pure fact that the filter coefficients are coded into the data stream; the manner of decoding may depend on the current block being a TNS or NS block) (Note that the generation of the filter parameters may depend on the current block being a TNS block or a NS block: in case of TNS, the filter parameters may be determined by LP (e.g. linear prediction) analysis of the transform coefficients or spectral coefficients to result into least error when comparing the TNS filtered version with the original version, and in case of NS with the aim of shaping flat pseudo noise in a pleasant way); wherein the filter parameters are to be used at the decoder for the TNS filtering, in case of the selected mode being the TNS mode; and the filter parameters are to be used for shaping, by the decoder, noise using which zero-quantized portions of the spectrum are to be filled at the decoder, in case of the selected mode being the NS mode.
[0044] According to an embodiment of the invention, (the encoder is optionally configured to selectively switch (e.g. a block-wise, e.g. frame-wise) between the TNS mode and the NS mode as the selected mode (e.g. as a current operating mode of the encoder) (e.g. between consecutive blocks)); wherein the encoder is configured to subject the spectrum of the current block to TNS filter analysis (e.g. in a sense of a analysis for a prediction filtering in a spectral, e.g. frequency, direction, e.g. over frequency, e.g. across frequency e.g. a spectral domain prediction filtering, e.g. in a sense of using an analysis for a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) to obtain TNS filter parameters; and apply the TNS filter parameters to the spectrum (e.g. of the current block) to obtain a TNS filtered spectrum; encode the TNS filtered spectrum into the data stream, and encode a filter information associated with the TNS filter parameters (e.g. a filter parametrization or a differential information based on which a filter parametrization may be determined or predicted), in case of the selected mode being the TNS mode; and wherein the encoder is configured to determine (e.g. to determine, e.g. to calculate) noise synthesis parameters on the basis of the spectrum of the current block, for subjecting the spectrum to NS in a (e.g. corresponding) decoder, encode the spectrum (e.g. of the current block), and encode a noise synthesis information associated with the noise synthesis parameters (e.g. the noise synthesis parameters or a differential information based on which the noise synthesis parameters may be determined or predicted), in case of the selected mode being the NS mode.
[0045] According to an embodiment of the invention, the encoder is configured to selectively switch (e.g. a block-wise, e.g. frame-wise) between the TNS mode and the NS mode as the selected mode (e.g. as a current operating mode of the encoder) (e.g. between consecutive blocks) (e.g. wherein the TNS mode and the NS mode are mutually exclusive); wherein the encoder is configured to subject the spectrum of the current block to TNS filtering (e.g. in a sense of a prediction filtering in a spectral, e.g. frequency, direction, e.g. over frequency, e.g. across frequency e.g. using a spectral domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) using TNS filter parameters, in order to obtain a TNS filtered spectrum (e.g. wherein the encoder is configured to subject a (e.g. any arbitrary) spectrum to TNS filtering, but wherein the encoder is optionally configured to subject the spectrum of the current block to TNS filtering using TNS filter parameters, in order to obtain a TNS filtered spectrum in case of the selected mode being the TNS mode); wherein the encoder is configured to obtain (e.g. to determine, e.g. to calculate) noise synthesis parameters on the basis of the spectrum of the current block, for subjecting the spectrum to NS in a (e.g. corresponding) decoder (e.g. wherein the encoder is configured to obtain noise synthesis parameters on the basis of a (e.g. any arbitrary) spectrum for subjecting the spectrum to NS, but wherein the encoder is optionally configured to obtain the noise synthesis parameters on the basis of the spectrum of the current block, for subjecting the spectrum to NS in a decoder, in case of the selected mode being the NS mode), wherein the encoder is configured to encode the TNS filtered spectrum (e.g. into the data stream), and encode a filter information associated with the TNS filter parameters (e.g. a filter parametrization or a differential information based on which a filter parametrization may be determined or predicted), in case of the selected mode being the TNS mode; and wherein the encoder is configured to encode the spectrum (e.g. of the current block), and encode a noise synthesis information associated with the noise synthesis parameters (e.g. the noise synthesis parameters or a differential information based on which the noise synthesis parameters may be determined or predicted), in case of the selected mode being the NS mode.
[0046] According to an embodiment of the invention, the encoder is configured to obtain (e.g. to determine, e.g. to calculate) the noise synthesis parameters as noise synthesis filter parameters, and wherein the noise synthesis information comprises an information about the noise synthesis filter parameters.
[0047] According to an embodiment of the invention, encoder is configured to predict the TNS filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream fora previous block (e.g. a previously encoded block) to obtain predicted TNS filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted TNS filter parameters and the TNS filter parameters for the current block), and encode the prediction residual (e.g. as the filter information or as a portion of the filter information); in case of the selected mode being the TNS mode; and wherein the encoder is configured to predict the noise synthesis filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted noise synthesis filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted noise synthesis filter parameters and the noise synthesis filter parameters for the current block), and encode the prediction residual (e.g. as the noise synthesis information or as a portion of the noise synthesis information); in case of the selected mode being the NS mode.
[0048] Here, it is to be noted that the encoder-sided prediction may comprise a reduction of redundant signal portions, which may be reconstructed on a decoder side, e.g. optionally based on prediction filter coefficients provided in the bitstream, e.g. in addition to the prediction residual. In general, the encoder-sided prediction may be considered an encoder-sided prediction processing and the decoder-sided prediction may be considered a decoder-sided prediction processing. These functionalities may, for example, be understood as inverse processings, e.g. reducing and respectively reconstructing redundant signal portions.
[0049] According to an embodiment of the invention, the encoder is configured to selectively switch (e.g. a block-wise, e.g. frame-wise) between a predictive mode (e.g. delta_time_flag having value 1) and a non-predictive mode (e.g. delta_time_flag having value 0) as a selected parameter mode; wherein the encoder is configured to predict the TNS filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for a previous block (e.g. a previously encoded block) to obtain predicted TNS filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted TNS filter parameters and the TNS filter parameters for the current block), and encode the prediction residual (e.g. as the filter information or as a portion of the filter information); in case of the selected mode being the TNS mode and in case of the selected parameter mode being the predictive mode; and wherein the encoder is configured to predict the noise synthesis filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted noise synthesis filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted noise synthesis filter parameters and the noise synthesis filter parameters for the current block), and encode the prediction residual (e.g. as the noise synthesis information or as a portion of the noise synthesis information); in case of the selected mode being the NS mode and in case of the selected parameter mode being the predictive mode; wherein the encoder is configured to encode the TNS filter parameters (e.g. in a non-differential manner, e.g. without predictive model or computing residuals) in the data stream for the current block independent from filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for the previous (e.g. previously encoded) block of the DWS, in case of the selected mode being the TNS mode and in case of the selected parameter mode being the non-predictive mode, wherein the encoder is configured to encode the noise synthesis filter parameters (e.g. in a non-differential manner, e.g. without predictive model or computing residuals) in the data stream for the current block independent from filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the datastream for the previous block (e.g. previously encoded) of the DWS, in case of the selected mode being the NS mode and in case of the selected parameter mode being the non-predictive mode; and wherein the encoder is configured to encode a parameter mode information (e.g. delta_time_flag) in the data stream, which indicates the selected parameter mode out of the predictive mode (e.g. delta_time_flag having value 1) and the non-predictive mode (e.g. delta_time_flag having value 0).
[0050] According to an embodiment of the invention, the parameter mode information (e.g. delta_time_flag) is a block-specific and / or channel specific information (e.g. wherein the DWS comprises a plurality of channels); and wherein the encoder is configured to selectively encode the TNS filter parameters and / or the noise synthesis filter parameters using the predictive mode or the non- predictive mode on a per-block basis and / or on a per-channel basis.
[0051] According to an embodiment of the invention, the encoder is configured to encode the TNS filter parameters and / or the noise synthesis filter parameters as reflection coefficients and / or as lattice coefficients or as prediction residuals thereof.
[0052] According to an embodiment of the invention, the encoder is configured to perform the prediction of the TNS filter parameters and / or of the noise synthesis filter parameters and / or the obtaining of the prediction residual (e.g. based on the predicted TNS filter parameters and the TNS filter parameters for the current block and / or based on the predicted noise synthesis filter parameters and the noise synthesis filter parameters for the current block) in a spectral domain (in other words, as an example, even prediction and / or correction may be done in this representation domain, e.g. corresponding to the reflection coefficients and / or as lattice coefficients).
[0053] According to an embodiment of the invention, the encoder is configured to predict the TNS filter parameters based on noise synthesis filter parameters signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted TNS filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted TNS filter parameters and the TNS filter parameters for the current block), and encode the prediction residual; in case of the selected mode being the TNS mode; and wherein the encoder is configured to predict the noise synthesis filter parameters based on TNS filter parameters signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted noise synthesis filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted noise synthesis filter parameters and the noise synthesis filter parameters for the current block), and encode the prediction residual; in case of the selected mode being the NS mode.
[0054] According to an embodiment of the invention, the encoder is configured to predict the TNS filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted TNS filter parameters, irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted TNS filter parameters and the TNS filter parameters for the current block), and encode the prediction residual; in case of the selected mode being the TNS mode; and wherein the encoder is configured to predict the noise synthesis filter parameters based on filter parameters (e.g. TNS filter parameters and / or noise synthesis filter parameters) signaled in the data stream for a previous block (e.g. a previously encoded block), to obtain predicted noise synthesis filter parameters, irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters, obtain a prediction residual (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters and the filter parameters for the current block) (e.g. based on the predicted noise synthesis filter parameters and the noise synthesis filter parameters for the current block), and encode the prediction residual; in case of the selected mode being the NS mode.
[0055] According to an embodiment of the invention, the encoder is configured to selectively switch (e.g. on a block-wise basis, e.g. on a frame-wise basis, e.g. on a channel-wise basis) between encoding TNS filter parameters and time-domain gain control parameters (e.g. with scalar multiplication of the decoded time-domain signal with transmission of the scalar values) (e.g. in order to subject the spectrum to TNS), in case of the selected mode being the TNS mode.
[0056] According to an embodiment of the invention, the encoder is configured to encode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode (e.g. TNS mode, e.g. NS mode, e.g. when having a value other than zero, e.g. a value of 2), which is to be considered (e.g. applied) for every block of the DWS.
[0057] According to an embodiment of the invention, the encoder is configured to encode a channelspecific high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode (e.g. the TNS mode, e.g. the NS mode, e.g. when having a value other than zero, e.g. a value of 2), which is to be considered (e.g. applied) for every block of a channel (or at least of a portion of this channel) of the DWS (e.g. wherein the DWS comprises a plurality of channels).
[0058] According to an embodiment of the invention, the encoder is configured to encode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is any out of the TNS mode or the NS mode (e.g. indicating whether the selected mode is a perceptual mode, e.g. comprising spectral shaping on the basis of perceptual characteristics); and wherein the encoder is configured to skip encoding the mode signal, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode (e.g. when percept_mode 2).
[0059] According to an embodiment of the invention, the mode signal comprises a m-ary syntax element (e.g. with m>=2), wherein a first value indicates that TNS mode and wherein a second value indicates the NS mode.
[0060] According to an embodiment of the invention, the encoder comprises an inventive decoder in order to predictively encode the DWS.
[0061] According to an embodiment of the invention, the encoder is configured to obtain a noise level information (e.g. a noise level energy), wherein the encoder is configured to obtain (e.g. to determine, e.g. to calculate) noise synthesis parameters as noise synthesis filter parameters, and wherein the noise synthesis information comprises an information about the noise synthesis filter parameters and the noise level information.
[0062] An embodiment according to the invention comprises a method for decoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) from a data stream using block-based transform decoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), the method comprising: decoding a spectrum (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum) for a current (e.g. a currently considered block, e.g. a currently encoded block) block of the DWS from the data stream, decoding a mode signal (e.g. percept_mode) from the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); subjecting the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value), in case of the selected mode being the TNS mode; and subjecting the spectrum to NS, in case of the selected mode being the NS mode.
[0063] An embodiment according to the invention comprises a method for encoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) into a data stream using block-based transform encoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), the method comprising: encoding a current block (e.g. a currently considered block, e.g. a currently encoded block) of the DWS into the data stream by deriving a spectrum of the current block (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum), encoding a mode signal (e.g. percept_mode) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the method comprises, in encoding the current block, encoding the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value), in case of the selected mode being the TNS mode; and subjecting the spectrum to NS, in case of the selected mode being the NS mode.
[0064] It is to be noted that the above-methods may be based on the same concepts, considerations and / or principles as the above-discussed encoders and decoders. Hence, the above-methods may comprise any of the above-discussed features, functionalities and details of respective decoders or encoders, e.g. in a corresponding or respectively adapted manner, both individually or taken in combination. An embodiment according to the invention comprises a computer program for performing any of the inventive methods, when the computer program runs on a computer.
[0065] An embodiment according to the invention comprises a data stream comprising: an encoded representation of a spectrum for a current block of a DWS, and an encoded representation of a mode signal (e.g. an encoded representation of a mode signal), which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode.
[0066] An embodiment according to the invention comprises a data stream having encoded therein a digital wave form signal using the method for encoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. aseismic, e.g. an acoustic, e.g. an audio signal) into a data stream using block-based transform encoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), the method comprising: encoding a current block (e.g. a currently considered block, e.g. a currently encoded block) of the DWS into the data stream by deriving a spectrum of the current block (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum), encoding a mode signal (e.g. percept_mode) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the method comprises, in encoding the current block, encoding the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value), in case of the selected mode being the TNS mode; and subjecting the spectrum to NS, in case of the selected mode being the NS mode.
[0067] It is to be noted that the above datastreams may be based on the same concepts, considerations and / or principles as the above-discussed encoders and decoders. Hence, the above datastreams may comprise any of the above-discussed features, functionalities and details of respective decoders or encoders, e.g. in a corresponding or respectively adapted manner, both individually or taken in combination.
[0068] Brief Description of the
[0069]
[0070] The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0071] Fig. 1 shows a schematic view of a decoder according to embodiments of the invention;
[0072] Fig. 2 shows a schematic view of a decoder with additional optional features, according to embodiments of the invention;
[0073] Fig. 3 shows a schematic view of an encoder according to embodiments;
[0074] Fig. 4 shows a schematic view of an encoder with additional, optional features, according to embodiments;
[0075] Fig. 5 shows a schematic view of an encoder for encoding a multi-channel digital signal into a datastream as well as a decoder for decoding the multi-channel digital signal from datastream according to embodiments; and
[0076] Fig. 6 shows a schematic view of a concept according to embodiments of the invention.
[0077] Detailed Description of the Embodiments
[0078] Equal or equivalent elements or elements with equal or equivalent functionality are denoted in the following description by equal or equivalent reference numerals even if occurring in different figures.
[0079] In the following description, a plurality of details is set forth to provide a more throughout explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present invention. In addition, features of the different embodiments described herein after may be combined with each other, unless specifically noted otherwise.
[0080] Fig. 1 shows a schematic view of a decoder according to embodiments of the invention. Fig. 1 shows decoder 100, for decoding a digital waveform signal, DWS, from a data stream using block-based transform decoding. Therefore, decoder 100 is provided with the data stream 101.
[0081] The decoder 100 is configured to, e.g. using decoding unit 110, decode a spectrum 111 for a current block of the DWS from the data stream 101 and to decode a mode signal 112, from the data stream 101, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode.
[0082] Furthermore, the decoder 100 is configured to subject the spectrum 111 to TNS filtering, e.g. using a TNS filtering unit 130, in case of the selected mode being the TNS mode and to subject the spectrum 111 to NS, e.g. using a NS unit 130, in case of the selected mode being the NS mode. The switching between the TNS filtering an NS processing, may be performed using a switching unit 120, which may be provided with the mode signal 112, in order to perform the switching.
[0083] Furthermore, as an optional feature, the decoder 100 may, for example, be configured to decode filter parameters 113 from the data stream 101 for the current block and to use the filter parameters for the TNS filtering, in case of the selected mode being the TNS mode and to fill zero-quantized portions of the spectrum using noise, and shape the noise using the filter parameters, in case of the selected mode being the NS mode.
[0084] Hence, optionally, for both the TNS filtering unit 130 and the NS unit 140, e.g. irrespective of whether the selected mode is the TNS mode or the NS mode, e.g. as indicated by the mode signal 112, only one instead of two sets of filter parameters 113 may be decoded, e.g. for a processing of the current block.
[0085] Next, reference is made to Fig. 2. Fig. 2 shows a schematic view of a decoder with additional optional features, according to embodiments of the invention. Decoder 200 is configured to obtain a data stream. In the example of Fig. 2 different optional contents of such a data stream are shown (e.g. 201 to 206, individually or in any combination). Hence, the data stream may comprise an encoded representation 201 of a spectrum to be decoded and optionally an encoded representation 202 of a mode information.
[0086] Decoder 200 may be configured to decode the encoded representation 201 of the spectrum, for example using a spectrum decoding unit 210, in order to obtain a decoded version of the spectrum 211.
[0087] The encoded representation 202 of the mode information may, for example, be decoded using a mode information decoding unit 210’. Based on the decoding, a mode signal 212 may be obtained. In other words, the mode information may comprise a mode signal. The mode signal 212 may optionally comprises a m-ary syntax element, e.g. with m>=2, wherein a first value indicates the TNS mode and wherein a second value indicates the NS mode.
[0088] Furthermore, decoder 200 comprises a TNS filtering unit 230, a NS unit 220 and a switching unit 220. Hence, in line with the decoder of Fig. 1, decoder 200 is configured to subject the spectrum 211 to TNS filtering, e.g. using TNS filtering unit 230, in case of the selected mode being the TNS mode (e.g. as indicated as selected mode by the mode signal 212, setting switching unit 220 accordingly) and to subject the spectrum 211 to NS, e.g. using NS unit 240, in case of the selected mode being the NS mode (e.g. as indicated as selected mode by the mode signal 212, setting switching unit 220 accordingly).
[0089] As an example, the NS unit 240 may be configured to fill zero-quantized portions of the spectrum 211 using noise and to shape the noise.
[0090] As an additional optional feature, decoder 200 comprises a transform unit 250, which may be configured to transform the TNS filtered spectrum 231 or the NS processed spectrum 241 (e.g. from a spectral domain to an original domain) for providing signal 251.
[0091] Furthermore, the decoder 200 may be configured, as an optional feature, to provide filter parameters 261 to the TNS filtering unit 230 and to the NS unit 240.
[0092] As an example, decoder 200 may be configured to decode the filter parameters 261 from the data stream, e.g. for the current block, and to use the filter parameters 261 for the TNS filtering, in case of the selected mode being the TNS mode and to use the filter parameters 261 to shape noise, introduced in zero-quantized portions of the spectrum 211, in case of the selected mode being the NS mode. Hence, the filter parameters 261 may be used for TNS or NS, e.g. in a mode agnostic manner.
[0093] Moreover, as an optional feature, the TNS filtering and a respective filtering performed for the NS filtering may comprise a same filter structure. Hence, decoder 200 may be configured to subject the spectrum 211 to TNS filtering, in case of the selected mode being the TNS mode and subject the spectrum 211 to NS, in case of the selected mode being the NS mode, using a same filter structure.
[0094] Decoder 200 may, for example, be configured to obtain the filter parameters 261 based on a predictive mode or based on a non-predictive mode, e.g. depending on how the filter parameters 261 were encoded on an encoder side.
[0095] Therefore, as an optional feature, the bitstream may comprise an encoded representation of a parameter mode information. The parameter mode information may indicate a selected parameter mode to be a predictive mode or a non-predictive mode. Here, as an example the parameter mode information 213 is included in the mode information (see 202). For example, a respective functionality may be switched within the decoder 200 using a switching unit 260.
[0096] For example, in case the filter parameters were encoded using an encoder-sided prediction processing, the bitstream may comprise an encoded representation 203 of a prediction residual of filter parameters, and as an optional feature, for example, in addition an encoded representation 204 of filter parameter prediction coefficients. Using respective decoding units (filter parameter residual decoding unit 210” and optionally filter parameter prediction coefficients decoding unit 210”’) the decoder may obtain a decoded version of the prediction residual of filter parameters 214 and optionally a decoded version of filter parameter prediction coefficients 215.
[0097] Hence, for example using a filter parameter prediction unit 270, the decoder may be configured to decode the (e.g. current) filter parameters 281 by to predicting the filter parameters based on filter parameters 28T signaled in the data stream for a previous block to obtain predicted filter parameters 272. Furthermore, the decoder may be configured to correct the predicted filter parameters 272 (e.g. using filter parameter correction unit 280) using the prediction residual 214 to decode the filter parameters 281.
[0098] Hence, the filter parameters 281 determined in a current coding step f for a current block may be, in coding step f+1, filter parameters 28T signaled in the data stream for a previous block (namely f) based on which the next prediction is performed. Hence, the loop of signal 281’ may symbolize a signal routing when proceeding from processing step f to processing step f+1, wherein current filter parameters 281 become previous filter parameters 28T.
[0099] In other words, the current filter parameters 281 may hence be the filter parameters based on which another prediction is performed or in other words in a next time step, the current filter parameters 281 may become the filter parameters 281’ signaled in the data stream for a previous block.
[0100] As an optional feature, in case an encoded representation 204 of filter parameter prediction coefficients are present in the bitstream, the decoded version thereof 215 may be used for the prediction of the filter parameters 272, e.g. as an input to unit 270.
[0101] Furthermore, as an optional feature, the above decoding of filter parameters 281 may be performed irrespective of whether the filter parameters 281 for the current block and the filter parameters 28T signaled in the data stream for the previous block are associated with a same selected mode, hence, e.g. irrespective of the mode signal 212 indicating, for step f, a same mode or a different mode, compared to mode signal 212 of step f-1.
[0102] Furthermore, as an optional feature, the decoder 200 may be configured to perform the prediction of the filter parameters and / or the correction of the predicted filter parameters in a spectral domain. Hence, units 270 and 280 may be configured to receive, process and provide spectral domain representations.
[0103] For example, in case the filter parameters were encoded using an encoder-sided processing without prediction, decoder 200 may be configured to decode the filter parameters 216 from the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS. Therefore, the data stream may comprise an encoded representation of filter parameters 205 which may be decoded using a decoding unit 210’”.
[0104] Hence, as discussed previously, in the data stream, an information about the encoder-sided processing mode for obtaining filter parameters 281 or 216 may be included, e.g. in the form of the parameter mode information 213. Hence, depending on the parameter mode information indicating the predictive or the non-predictive mode, the decoder 200 may obtain and process filter parameters 281 or 216, e.g. as indicated by switching unit 260. Optionally, the parameter mode information 213 may be a block-specific and / or channel specific information. Hence, as an example, the decoder may be configured to selectively decode the filter parameters using the predictive mode or the non-predictive mode on a perblock basis and / or on a per-channel basis.
[0105] For example, the decoder 200 may be configured to decode the filter parameters 261 as reflection coefficients and / or as lattice coefficients.
[0106] As another optional example, the decoder 200 may be configured to obtain an encoded representation 206 of noise synthesis side-information (e.g. comprising noise level and spectral shape parameters; e.g. comprising an information about a noise level energy and a spectral tilt, e.g. noise synthesis parameters which are not filter parameters). The decoder may be configured to decode the same using noise synthesis side-information decoding unit 210”” and to provide the decoded version of the noise synthesis side-information 217 to the NS unit 240.
[0107] As an example, the NS unit 240 may be configured to fill zero-quantized portions of the spectrum 211 using noise, and shape the noise using the noise synthesis side-information 217 and / or using filter parameters, e.g. 216 or 281, in case of the selected mode being the NS mode.
[0108] Here, it is to be noted that decoder 200 may be configured to achieve a noise filling and shaping functionality, for example, independent of the previously discussed filtering using filter parameters 261, e.g. based solely on the noise synthesis side-information. For example, in this case the side-information 217 may not be considered a side-information (e.g. for use in combination with filter parameters), but a noise synthesis information or for example noise synthesis parameters, e.g. to achieve an independent noise synthesis functionality. Hence, accordingly, the decoder 200 may be configured to switch between respective operating modes, e.g. TNS vs. NS, e.g. predictive vs. non-predictive, e.g. filter-parameter based NS or non-filter-parameter based NS, e.g. TNS using spectral domain filtering vs. TNS using time domain gain control.
[0109] Hence, as an example, the encoded representation 206 of noise synthesis side-information may as well comprise a supplementary information for the filtering-based noise shaping in unit 240. For example, the noise synthesis parameters 206 may comprise a noise level information, so that unit 240 (given the selected mode is the NS mode) may fill zero-quantized portions of the spectrum 211 using noise and shape the noise using the filter parameters 261 and using the noise level information.
[0110] As discussed above, in case of the selected mode being the TNS mode, the decoder 200 may use said filter parameters for the TNS filtering of the spectrum 211.
[0111] Furthermore, in view of the TNS functionality, the decoder 200 may be configured to perform, e.g. as an alternative the previously discussed TNS filtering using filter parameters 261, timedomain gain control in order to perform TNS. Hence, optionally, the decoder 200 may be configured to selectively switch between decoding filter parameters and using the filter parameters for the TNS filtering, and performing time-domain gain control in order to perform TNS in case of the selected mode being the TNS mode. Accordingly, decoder 200 may be configured to obtain a decoded version of such time-domain gain control parameters and perform a respective switching.
[0112] Furthermore, as an optional feature, the mode information may comprise a high level syntax element, e.g. percept_mode, e.g. in the form of a syntax flag, indicating whether the selected mode is a perceptual mode which is to be considered for every block of the DWS.
[0113] Furthermore, as an optional feature, the mode information may comprise a channel-specific high level syntax element indicating whether the selected mode is a perceptual mode for every block of a channel of the DWS.
[0114] Furthermore, as an optional feature, the mode information may comprise a high level syntax element indicating whether the selected mode is any out of the TNS mode or the NS mode based on which the decoder 200 may selectively decide to skip decoding the mode signal 212, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode.
[0115] Next reference is made to Fig. 3. Fig. 3 shows a schematic view of an encoder according to embodiments. Encoder 300 is an encoder for encoding a digital waveform signal 301, DWS, into a data stream using block-based transform encoding. The encoder 300 is configured to encode a current block of the DWS into the data stream by deriving a spectrum 311 of the current block (e.g. using transform unit 310), encode a mode signal 321 (which may, optionally, be obtained by a mode signal determination unit 320 and encoded using mode signal encoding unit 330) into the data stream (e.g. as encoded representation 331 of the mode signal), which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode.
[0116] However, optionally, the mode signal 321 may as well be received from an external source.
[0117] Furthermore, the encoder 300 is configured to, in encoding the current block, encode the current block so that the current block is to be decoded from the data stream by subjecting the spectrum 311 to TNS filtering, in case of the selected mode being the TNS mode; and subjecting the spectrum 311 to NS, in case of the selected mode being the NS mode.
[0118] Here, as optional features, the encoder 300 may be configured, e.g. in case of the selected mode being the TNS mode, to determine, based on spectrum 311, TNS filter parameters 341, using TNS filter analysis unit 340.
[0119] Based on said TNS filter parameters 341, an encoder-sided TNS filtering of the spectrum 311 may be performed, in order to encode the TNS filtered spectrum 351, using spectrum encoding unit 360. Optionally, the TNS filter parameters 341 (or a respective information to reconstruct said filter parameters on a decoder side, e.g. a residual information, e.g. additionally, with filter parameter prediction coefficients) may be encoded using processing information encoding unit 370, in order to obtain an encoded representation of processing information 371.
[0120] Here, as optional features, the encoder 300 may be configured, e.g. in case of the selected mode being the NS mode, to determine, based on spectrum 311 , a noise synthesis information 381, using NS parameter determination unit 380. In this case, the noise synthesis information may be encoded as the processing information (see 371) and encoded representation 361 may be provided based on the unfiltered spectrum 311 (e.g. wherein some spectral portions may be quantized to zero for a decoder-sided noise filling and shaping based on encoded noise synthesis information 381).
[0121] Hence, the processing information may comprise TNS filter parameters 341 or noise synthesis information 381. Optionally, the processing information may, for example, comprise one or more of the previously discussed filter parameter prediction coefficients and / or prediction residual of filter parameters (e.g. as a representation of the TNS filter parameters 341), the filter parameters themselves, e.g. for TNS or NS (e.g. given information 381 comprises filter parameters) and / or (e.g. in addition) a noise synthesis side-information, e.g. depending on the operating mode) The different operating modes and a respective switching thereof are indicated using switching units 390, wherein the switching may be performed based on the mode signal 321 .
[0122] For example, given the NS mode, the spectrum encoding unit 360 may be configured to quantize some portions of spectrum 311 to zero for a subsequent decoder-sided noise filling.
[0123] Next, reference is made to Fig. 4. Fig. 4 shows a schematic view of an encoder with additional, optional features, according to embodiments.
[0124] Encoder 400 is an encoder for encoding a digital waveform signal 401, DWS, into a data stream using block-based transform encoding. Therefore, the encoder 400 is configured to encode a current block of the DWS into the data stream by deriving a spectrum 411 (using transform unit 410) of the current block, and to encode a mode signal 421 (e.g. determined by an optional mode signal determination unit 420, e.g. based on an analysis of spectrum 411 ; e.g. alternatively received from an external source) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode. Here as an example, an encoded representation 402 of a mode information is encoded in the bitstream, using an optional mode information encoding unit 430, the mode information comprising the mode signal 421. Hence, the mode information may be the mode signal 421 or may comprise additional information, e.g. side-information, e.g. such a parameter mode information.
[0125] Furthermore, the encoder 400 is configured to, in encoding the current block, encode the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering, in case of the selected mode being the TNS mode; and subjecting the spectrum to NS, in case of the selected mode being the NS mode.
[0126] Therefore, as an optional feature, encoder 400 is configured to encode filter parameters into the data stream for the current block for use with a decoder-sided TNS filtering, in case of the selected mode, as indicated by mode signal 421, being the TNS mode or for use with a decoder-sided noise shaping of zero-quantized portions of the spectrum 411, in case of the selected mode being the NS mode.
[0127] Therefore, encoder 400 may, for example, be configured to provide an encoded representation 403 of a spectrum (e.g. using optional spectrum encoding unit 430’) on the basis of an encodersided TNS filtered spectrum 451 (e.g. using the optional encoder-sided TNS filtering unit 450) or on the basis of the spectrum 411 (e.g. wherein some spectral values may be quantized to zero for a subsequent decoder-sided noise filling and shaping), depending on the selected mode, as indicated by switch 490.
[0128] Accordingly, the encoder 400 may be configured to determine TNS filter parameters 441 for TNS filtering, e.g. using TNS filter analysis unit 400, for example based on the spectrum 411 . In other words, the encoder may be configured to subject the spectrum 411 of the current block to TNS filter analysis to obtain TNS filter parameters 441. An information about said filter parameters may, optionally, be provided to the filtering unit 450 for obtaining the TNS filtered spectrum 451. The encoder may be configured to apply the TNS filter parameters 441 to the spectrum 411 to obtain a TNS filtered spectrum and to encode the TNS filtered spectrum 451 into the data stream, in case of the selected mode being the TNS mode.
[0129] Here, it is to be noted that for the sake of simplicity encoder-sided TNS filter parameters and decoder-sided TNS filter parameters may be commonly referred to as TNS filter parameters. However, it should be noted that optionally respective versions of said filter parameters adapted to the respective side (e.g. decoder-side, e.g. encoder-side) may be obtained and used and respectively transmitted. Hence, for example the encoder may obtain TNS filter parameters 341 ,441 for use with an encoder-sided TNS filtering and may convert the same to be adapted to a decoder-sided TNS filtering (e.g. an inverse filtering compared to the encodersided filtering) before encoding the same. However, such optional, implicit conversion and adaptation steps are not shown in Fig. 3 and 4. Furthermore, optionally, the encoder may transmit the same parameters used encoder-sided and a respective decoder may perform any potential adaptations. The same applies to NS filter parameters accordingly.
[0130] Accordingly, encoder 400 may be configured to determine noise synthesis filter parameters 481, e.g. using NS parameter determination unit 480, for use in a decoder-sided noise filling and / or shaping. Furthermore, the encoder 400 may be configured to determine a noise synthesis side-information 482 (e.g. comprising one or more of a noise level, spectral shape parameters, e.g. a spectral tilt information), for use with a decoder-sided noise filling and / or shaping.
[0131] Here, it is to be noted that the noise synthesis filter parameters 481 in combination with the noise synthesis side-information 482 may also be considered noise synthesis parameters or a noise synthesis information, e.g. noise synthesis parameters or a noise synthesis information obtained on the basis of the spectrum 411 of the current block, for subjecting the spectrum to NS in a decoder (e.g. in case of the selected mode being the NS mode). Noise synthesis parameters may hence optionally comprise or even be filter parameters. Accordingly, such a noise synthesis information or parameters may be encoded in the data stream, in case of the selected mode being the NS mode.
[0132] However, again, optionally, information 482 may be an independent noise synthesis information, for performing a decoder-sided noise filling and shaping, which is not a filtering. Hence, information 405 may as well be provided without any filter information 404.
[0133] As an optional feature, e.g. depending on the selected mode, as indicated by mode signal 421, based on the TNS filter parameters 441 or based on the noise syntheses filter parameters 481 an information about filter parameters may be encoded in the bitstream (as will be discussed in detail hereafter).
[0134] Hence, as an example, encoder 400 may be encoder is configured to selectively switch, see switch 520, between the TNS mode and the NS mode as the selected mode. Hence, as discussed above, the encoder 400 may be configured to subject the spectrum 411 of the current block to TNS filtering, 450, using TNS filter parameters, 441, in order to obtain a TNS filtered spectrum 451; and to obtain noise synthesis parameters (e.g. 481 and / or 482) on the basis of the spectrum, 411 , of the current block, for subjecting the spectrum to NS in a decoder
[0135] Accordingly, the encoder may be configured to encode the TNS filtered spectrum 451, and encode a filter information (e.g. at least one of 407, 406, 404, e.g. 406, e.g. 406+407, e.g. 404) associated with the TNS filter parameters, in case of the selected mode being the TNS mode; and to encode the spectrum 411, or at least portions thereof (e.g. when some spectral values are quantized to zero in unit 430’), and to encode a noise synthesis information associated with the noise synthesis parameters (e.g. at least one of 407, 406, 404, 405, e.g. 406, e.g.
[0136] 406+407, e.g. 404, e.g. 404+405) in case of the selected mode being the NS mode.
[0137] As another optional feature, the encoder 400 may be configured to selectively switch between a predictive mode and a non-predictive mode as a selected parameter mode (e.g. to decide on how to encode said information about the filter parameters). Optionally, the encoder 400 may determine a parameter mode information 42T, e.g. using an optional parameter mode determination unit 420’, indicating or setting the selected parameter mode. Alternatively, encoder 400 may be provided with such an information from an external source.
[0138] Optionally, the mode information may comprise the parameter mode information 42T and may hence be provided to a respective decoder. As an optional feature, the encoder 400 may be configured to selectively switch between encoding TNS filter parameters 461 and time-domain gain control parameters, in case of the selected mode being the TNS mode. Hence, the encoder 400 may optionally be configured to obtain such time-domain gain control parameters, e.g. using unit 440, which may optionally, be provided with a time-domain version of the spectrum 441, e.g. with signal 401 directly, to obtain the parameters. Accordingly, encoder 400 may comprise another switching unit for providing TNS filter parameters 441 as filter coefficients or to provide time-domain gain control parameters in the TNS mode. Accordingly, TNS filtering in unit 450 may be performed differently.
[0139] Furthermore, as an optional feature, the encoder 400 may be configured to encode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode, which is to be considered for every block of the DWS. As an example, the parameter mode information 42T may comprise such a high-level syntax element.
[0140] Optionally, the encoder 400 may be configured to encode a channel-specific high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is a perceptual mode, which is to be considered for every block of a channel of the DWS.
[0141] As an example, the parameter mode information 42T may comprise such a channel-specific high-level syntax element.
[0142] As an example, the parameter mode information 42T may hence be a channel-specific information or for example a block-specific information. Accordingly, the encoder 400 may be configured to selectively encode the TNS filter parameters and / or the noise synthesis filter parameters using the predictive mode or the non- predictive mode on a per-block basis and / or on a per-channel basis
[0143] Optionally, the encoder 400 may be configured to encode a high level syntax element (e.g. percept_mode, e.g. a syntax flag), indicating whether the selected mode is any out of the TNS mode or the NS mode and to skip encoding the mode signal, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode. As an example, the mode information may comprise such a high level syntax element.
[0144] For example, the mode signal 421 may comprise a m-ary syntax element, wherein a first value indicates that TNS mode and wherein a second value indicates the NS mode. For example, given the non-predictive parameter mode, filter parameters 461, for a current block (e.g. TNS filter parameters 441 or noise synthesis filter parameters 481 depending on the selected mode as indicated by mode signal 421) may optionally be encoded independent from filter parameters of the previously encoded block (e.g. using filter parameter encoding unit 430” to provide a respective encoded representation 404).
[0145] For example, in case of the selected mode being the NS mode, in addition to the encoded representation of the filter parameters (e.g. irrespective of being encoded in a predictive or non-predictive mode), an encoded representation 405 of noise synthesis side-information may be provided (e.g. using encoding unit 430”’), e.g. comprising an information (e.g. additional information, e.g. side-information) for a respective decoder-sided noise filling. Hence, for example, a combination of signals 404 and 405 may be considered an encoded representation of noise synthesis parameters or an encoded representation of a noise synthesis information. However, as discussed previously, noise synthesis side-information 482 may optionally, as well be provided in a NS mode without providing filter parameters, e.g. for a decoder-sided noise filling and shaping that is not based on filtering.
[0146] Hence, in general, encoder 400 may be configured to obtain a noise level information, e.g. a noise level energy, e.g. using unit 480 and as a part of the side information 482, to obtain noise synthesis parameters as noise synthesis filter parameters 481, and to encode a noise synthesis information, e.g. signals 404+405, so that the noise synthesis information comprises an information about the noise synthesis filter parameters and the noise level information.
[0147] As another optional feature, e.g. in case of the selected parameter mode (e.g. as indicated by parameter mode information 421’) being the predictive mode (see switch 460), the filter parameters 461 may be encoded using an encoder-sided prediction processing.
[0148] Therefore, based on the current filter parameters 461 and predicted filter parameters 471, e.g. using an optional filter parameter residual determination unit 470, a prediction residual of filter parameters 472 may be obtained and encoded, e.g. using encoding unit 430”” to provide an encoded representation 406.
[0149] Optionally, encoder 400 may be configured to obtain the predicted filter parameters 471 , based on filter parameters 521 from a previous encoding step (e.g. using a filter parameter prediction unit 520). The prediction may, optionally, be performed using a set of prediction coefficients 522, hence, optionally, an information about the prediction coefficients 522 for predicting the TNS / NS filter parameters 471, may, for example, be encoded in the bitstream, e.g. using encoding unit 430””’, e.g. to obtain encoded representation 407.
[0150] As an optional feature, for obtaining the predicted filter parameters 471 , used for obtaining the current residual 472, the decoder-sided reconstruction of the filter parameters may be simulated encoder-sided. Hence, as an example, using a feedback loop, predicted filter parameters 471’ for a previous processing step (e.g. which were determined as the current, predicted filter parameters 471 but for the preceding block of signal 401 , e.g. in step f-1) may be corrected using a decoded version 501 of the prediction residual 406 for the previous processing step (e.g. which were determined for the preceding block of signal 401 , e.g. in step f-1). Hence, as optional features, encoder 400 comprises a decoding unit 500 and a correction unit 510.
[0151] Hence, the encoding may be performed on a block-by-block basis, wherein based on previously encoded residuals 501 (e.g. an encoded and subsequently decoded version thereof, e.g. from step f-1) and previously obtained predicted filter parameters 471’, the encoder may simulate the filter parameters 471, a respective decoder may predict for the current block, e.g. in step f.
[0152] The encoder may hence perform a prediction 520 alike to the decoder, to obtain new predicted filter parameters 471 for the current block, e.g. now in step f. Hence, again, the loop of signal 47 T may represent a jump from processing step f-1 to f, so that current predicted filter parameters (e.g. for a current block), 471, become predicted filter parameters 47T for the previously coded block. Now having the information that the decoder may obtain, the difference to the desired filter parameters 461 may be determined as the prediction residual 472 of the current block and provided as signal 406 to the decoder.
[0153] Hence, encoder 400 may optionally, comprise a decoder according to Fig. 3 or Fig. 4, in order to predictively encode the DWS.
[0154] As an example, the encoder 400 may perform the prediction of the TNS filter parameters 441 and / or of the noise synthesis filter parameters 521 and / or the obtaining of the prediction residual in a spectral domain.
[0155] Moreover, as an optional feature, encoder 400 may be configured to encode the TNS filter parameters and / or the noise synthesis filter parameters (see e.g. 404) as reflection coefficients and / or as lattice coefficients or as prediction residuals, see e.g. 406, thereof. Regarding the embodiment of Fig. 4, it is to be noted that filter parameters used for TNS may be predicted out of filter parameters used for NS and vice versa.
[0156] Hence, in other words, the encoder 400 may be configured to predict the TNS filter parameters (e.g. 461 in the TNS mode) based on filter parameters signaled in the data stream for a previous block (e.g. as reconstructed in signal 521), to obtain predicted TNS filter parameters (e.g. 471 in the TNS mode), irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters in case of the selected mode being the TNS mode; and to predict the noise synthesis filter parameters (e.g. 461 in the NS mode) based on filter parameters signaled in the data stream for a previous block (e.g. as reconstructed in signal 521), to obtain predicted noise synthesis filter parameters (e.g. 471 in the NS mode), irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters in case of the selected mode being the NS mode. In both cases a respective residual 472 (e.g. a residual spectral domain representation, e.g. a residual between predicted filter parameters 471 and the filter parameters for the current block 461 may be determined and encoded.
[0157] Furthermore, regarding the predictive mode or the non-predictive mode, the encoder 400 may be configured to selectively switch between these modes. As indicated in Fig. 4, the encoder 400 may obtain a respective parameter mode information 42T on its own, or may optionally be provided with such an information externally.
[0158] Therefore, as an optionally feature, a switch 460 is shown in Fig. 4. Hence, as a general aspect, the encoder 400 may be configured to selectively switch between a predictive mode and a non-predictive mode as a selected parameter mode, for example on the basis of signal 42T and using switch 460.
[0159] Hence, the encoder 400 may be configured to predict the TNS filter parameters (e.g. 461 in case of the selected mode being the TNS mode) based on filter parameters signaled in the data stream for a previous block (e.g. as reconstructed by signal 521) to obtain predicted TNS filter parameters (e.g. 471 in case of the selected mode being the TNS mode), obtain a prediction residual 472, and encode the prediction residual, see 406, in case of the selected mode being the TNS mode and in case of the selected parameter mode being the predictive mode. Furthermore, the encoder may be configured to predict the noise synthesis filter parameters (e.g. 461 in case of the selected mode being the NS mode) based on filter parameters signaled in the data stream for a previous block (e.g. as reconstructed by signal 521), to obtain predicted noise synthesis filter parameters (e.g. 471 in case of the selected mode being the NS mode), obtain a prediction residual 472, and encode the prediction residual, see 406 in case of the selected mode being the NS mode and in case of the selected parameter mode being the predictive mode.
[0160] The above processing may be performed in case switch 460 is switched to its upper output.
[0161] Furthermore, the encoder 400 may be configured to encode the TNS filter parameters (in the TNS mode) and respectively the noise synthesis filter parameters (in the NS mode), see e.g. signal 461 , in the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS, in case of the selected parameter mode being the non-predictive mode.
[0162] The above processing may be performed in case switch 460 is switched to its lower output.
[0163] Furthermore, accordingly, the encoder 400 may be configured to encode a parameter mode information 42T in the data stream, which indicates the selected parameter mode out of the predictive mode and the non-predictive mode.
[0164] In the following, different inventive embodiments and aspects will be described in sections “Efficient Signaling of Noise Synthesis and Shaping Parameters in General-Purpose Perceptual Waveform Coding”, in particular in the sections “Introduction to embodiments”, “Motivation according to embodiments”, “Summary of aspects of embodiments according to the Invention”, “Preferred Embodiment of Aspect 1” and “Preferred Embodiment of Aspect 2”.
[0165] Also, further embodiments will be defined by the enclosed claims.
[0166] It should be noted that any embodiments as defined by the claims and the preceding disclosure can be supplemented by any of the details (features and functionalities) described in the above-mentioned sections.
[0167] Also, the embodiments described in the above-mentioned sections can be used individually, and can also be supplemented by any of the features in another section, or by any feature included in the claims or in the preceding disclosure. Also, it should be noted that individual aspects described herein can be used individually or in combination. Thus, details can be added to each of said individual aspects without adding details to another one of said aspects.
[0168] Moreover, features and functionalities disclosed herein relating to a method can also be used in an apparatus (configured to perform such functionality). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses.
[0169] Moreover, features and functionalities disclosed herein relating to a method, in particular an encoding method can also be used in a data stream or bitstream (e.g. defining a respective data stream or bitstream element). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding data stream, e.g. as a resulting data stream as providing by said encoder. In other words, the data streams disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses and methods.
[0170] Also, any of the features and functionalities described herein can be implemented in hardware or in software, or using a combination of hardware and software, as will be described in the section “Implementation alternatives”.
[0171] The following section may be referred to under the title: “Efficient Signaling of Noise Synthesis and Shaping Parameters in General-Purpose Perceptual Waveform Coding”
[0172] The following may hence address in particular features, functionalities and details regarding embodiments for an efficient signaling of noise synthesis and shaping parameters in general-purpose perceptual waveform coding.
[0173] The following subsection may be referred to by the title “Introduction to embodiments”.
[0174] Hence, embodiment discussed herein will be introduced, for example, with respect to underlying techniques on which embodiments improve upon.
[0175] General-purpose perceptual coding of digital waveform signals, such as biophysical (e.g., medical), geophysical (e.g., seismic), or acoustic (e.g., audio) signals comprises, or for example even requires, three types of basic functionality: 1. spectral shaping of the quantization distortion, e.g. most commonly or optionally by filtering of the (re)quantized decoded signal coefficients, optionally, with a linear predictive coding (LPC)-like, for example spectrally "non-flat" filter; this functional component is often referred to as "spectral noise shaping", abbreviated SNS here,
[0176] 2. temporal shaping of the quantization distortion, e.g. eithervia time-domain gain control with scalar multiplication of the decoded time-domain signal with transmission of the scalar values, or e.g. via temporal noise shaping (TNS) of transformed versions of the decoded signal coefficients; see [1],
[0177] 3. synthesis of noise-like components (for example, most commonly background or recording noise) missing in the decoded signals, e.g., due to coarse quantization of said transformdomain signal coefficients; [2],
[0178] The SNS functional component is well studied and may be based on or may for example simply require time-domain synthesis filtering of the decoded (i.e. , frame-wise reconstructed) waveform signal, for example preferably with time-varying filters.
[0179] Embodiments may comprise any or all of the above basic functionalities, in combination or individually. In particular, embodiments may combine a mutually exclusive application of TNS filtering and noise synthesis, in addition to a selective application of SNS. In particular, embodiments may comprise a selective switching between any of the above-discussed approaches for TNS or NS, and approaches as discussed in the following.
[0180] The following subsection may be referred to by the title “Motivation according to embodiments”. Hence, problems, which embodiments according to the invention may solve will be discussed.
[0181] At least some of the TNS and noise synthesis (NS) components may, for example, be known from prior art, with the latter aspect also being referred to as "noise filling" in [2], but the following particularities are worth being noted:
[0182] - The inventors recognized that in all known prior art, both TNS and NS may be applied during the decoding of a frame, despite the parameter signaling overhead this causes; 2 parameter sets, one for each aspect, are needed (e.g. according to conventional approaches),
[0183] (Accordingly, embodiments comprise a mutually exclusive indication of whether a TNS or an NS mode is to be applied, so as to limit the amount of parameter sets necessary) - The inventors recognized that in all known prior art, similarities between the main parameters of either TNS or NS in successive (i.e. temporally adjacent) frames f of a certain waveform signal channel c are not being exploited, (Accordingly, embodiments according to the invention comprise, as optional features, an encoding and respectively decoding of parameters for TNS based on parameters for NS and vice versa. Hence, optionally, based on TNS / NS parameters of a preceding block or frame (e.g. in encoding / decoding order), parameters for NS / TNS (or TNS / NS) for a current block or frame may be predicted and corrected using a residual information. Hence, embodiments optionally comprise a prediction of parameters for NS or TNS, irrespective of whether NS or TNS parameters form a basis for said prediction)
[0184] - The inventors recognized that in low-energy, e.g. non-transient frames, NS may, perceptually, be more important than TNS support when transform coding is used, for example since large spectral parts of f in c may be or will be sparse (e.g. quantized to 0); For example without NS, this may often lead to low-pass blurring / muffling effects in the timedomain signal for f,
[0185] - The inventors recognized that in higher-energy, e.g. often transient frames, TNS may, for example, be perceptually more important than NS support in some or even most cases, for example since the transform coding may not or will not lead to a sparse time-domain signal for f but may, for example, produce temporal ringing artefacts, e.g. when a transient signal "spike" occurs in f, reducible (e.g. which may be reducible) via TNS.
[0186] (Accordingly, a mode signal may be determined or chosen (e.g. a selected mode may be set) based on the signal to be encoded, in particular based on a “transience-characteristic” of the signal or a portion thereof. Such information may, for example, be frame-specific, block-specific and / or channel specific)
[0187] In the following, two modifications according to embodiments to the above-mentioned solutions, (e.g. at least partially or e.g. partially) known from the prior art, are described, addressing, for example inter alia, said drawbacks of parameter signaling overhead and / or parameter redundancy.
[0188] The following subsection may be referred to by the title “Summary of aspects of embodiments according to the Invention”. Hence, some aspects of embodiments and in particular some preferred features of embodiments or respectively embodiments will be discussed:
[0189] Inter alia, to address said prior-art shortcomings, the following two modification aspects are being proposed (embodiments may comprise respective features, functionalities and details both individually and taken in combination):
[0190] 1. mutually exclusive signaling of TNS or NS functionality, optionally for a given f in c, for example, for a low parameter rate,
[0191] 2. harmonization and / or delta-time signaling (e.g. harmonization & delta-time signaling) of the TNS and NS parameters, for example requiring most of the rate. As an example, a coding concept according to embodiments may allow a selective switching between predicting TNS parameters and NS parameters (e.g. respective filter parameters), e.g. irrespective of previously predicted (e.g. temporally previous, e.g. parameters related to a previously decoded frame or channel or frame in a channel, or in general block) parameters being TNS parameters or NS parameters (in particular respective filter parameters).
[0192] Each of these two aspects will be discussed separately hereafter along with preferred embodiments. However, it is to be noted that both aspects may be used independently as well as in combination.
[0193] Preferred Embodiment of Aspect 1
[0194] Let percept_mode be a syntax flag, preferably (hence optionally) transmitted once per c or entire signal and indicating that either TNS or NS parameters (e.g. as mutually exclusive options) are present, for example, in each f in c (for example when percept_mode has a certain value other than zero). When that is the case (i.e., percept_mode holds that certain value, preferably 2), let block_percept_mode be another (e.g. binary) syntax flag, for example, transmitted once per f and c and indicating, e.g. in said f, whether TNS (preferably, via value 1) or NS parameters (preferably, via value 0) are present. In the other case (i.e., when percept_mode 2, as an optional example), block_percept_mode may, for example not be transmitted and may hence, for example, in particular not be transmitted in any f of c, and neither TNS nor NS may, for example, be applied, implying, for example, that neither TNS nor NS parameters are signaled.
[0195] Following prior-art embodiments, frames for which NS functionality has been signaled (preferably, via block_percept_mode = 0), may or shall transmit noise synthesis parameters , e.g. the required noise synthesis parameters, most notably and preferably, as an example, noise level (e.g. energy) and / or spectral shape (e.g. tilt or for example rather "envelope",) parameters, for example, associated with said f in c. Likewise, frames for which TNS functionality has been signaled (preferably, via block_percept_mode = 1), may, for example, or shall transmit, optionally conventional, temporal noise shaping parameters, most notably and preferably, for example, a spectral application range (e.g. length) parameter and / or filter coefficients (e.g. envelope) associated with fin c.
[0196] As an example, according to embodiments, if a high level syntax element has a predetermined value, e.g. percept_mode=2, a mode signal, e.g. block_percept_mode, may be encoded / decoded, indicating whether NS or TNS is to be used (and indicating that respective parameters are present in a respective data stream). For example, in the NS mode, a noise synthesis information comprising noise synthesis parameters may be encoded / decoded, as an example, comprising noise level (e.g. energy) and spectral shape (e.g. tilt) parameters. For example, in the TNS mode, an information may be encoded / decoded comprising, for example, a spectral application range (e.g. length) parameter and filter coefficients (e.g. envelope).
[0197] Furthermore, as an example, in the NS mode, the noise synthesis information may alternatively comprise noise synthesis filter parameters and in the TNS mode, a filter information comprising TNS filter parameters may be encoded / decoded.
[0198] Preferred Embodiment of Aspect 2
[0199] TNS envelopes are, for example, or even typically, represented by quantized lattice or reflection coefficients, for example specifying an order-L infinite (FIR) or finite (HR) impulse response filter, for example, for application during decoding. NS shape parameters, however, are, for example, or even usually, represented in a different (e.g., tilt) domain. To harmonize the two representation, according to embodiments, optionally, use of a (FIR or HR) lattice and / or reflection coefficient filter domain also for NS spectral shape (e.g., tilt) parametrization is proposed (embodiments hence optionally comprising NS parameters in said domains). Usually and preferably, for example, a low filter order L, optionally less than or equal to the order used for TNS filter signaling may suffice in this case. Moreover, since the TNS or NS filter coefficients may, for example, take up most of the side information rate needed for signaling of such techniques, time differential coding of the (preferably quantized) filter coefficients optionally across adjacent frames fin c, for example regardless of whether a given f has TNS or NS parameters associated therewith, is proposed and may hence optionally be used according to embodiments; Fig. 6. To this end, a further (e.g. binary) syntax flag, delta_time_flag (e.g. as an example for a parameter mode information) may, for example, be introduced, optionally in each f and c for which the associated percept_mode (e.g. an example for a high level syntax element) holds said certain value (preferably 2) indicating the presence of either TNS or NS parameters. This delta_time_flag may indicate that each filter coefficient is (e.g. value 1) or is not (e.g. value 0) transmitted differentially to the respective TNS or NS filter coefficient, for example, in the previous of c.
[0200] Note that, above, f may denote not only a frame instance but also, more generally, a block instance. For example, a block may correspond to a frame, or as another example, a frame may comprise a sequence of blocks (e.g. temporally adjacent blocks). For example, a block may correspond to a window, e.g. a window of a frame, or a block may, for example, describe a channel or a portion of a channel within one frame. The above description is extended in the following by the presentation of further embodiments. Before this, however, the description proceeds with a presentation of a possible framework or codec into which the embodiments described above as well as the embodiments described further below may be built into. Many details described in this framework are, however, optional when being combined with any of the above or subsequently described embodiments. To be more precise, the framework is described with respect to Fig. 5 which shows an encoder for encoding a multi-channel digital signal 14 into a datastream 16 as well as decoder 12 for decoding the multi-channel digital signal 14 from datastream 16. This description of Fig. 5 shall be seen as a presentation of new embodiments of the present application which result when combining any of the embodiments described above or any of the embodiments described subsequently is combined with the decoder 12 or encoder 10 of Fig. 5 either by adopting all details / functionalities described with respect to Fig. 5 or with leaving-out some of the details / functionalities described with respect to Fig. 5. Sometimes such “optional” features of Fig. 5 are explicitly identified as being optional with respect to the combination of the previously and subsequently described embodiments, but the just-mentioned possible combinations of the previously / subsequently explained embodiments with the description of Fig. 5 shall not be restricted to the these explicitly identified variations of Fig. 5 in terms of leaving-out certain features.
[0201] In Fig. 5, the multi-channel digital signal 14 is illustrated by way of an array of samples with the samples being illustrated as small squares 18. Each line / row corresponds to a certain channel of the multi-channel digital signal 14. Each channel of signal 14may have associated therewith a respective channel ID and Fig. 5 shows these channels as being ordered according to their channel ID along vertical axis 20 which, thus, corresponds to a “source” channel axis 20. The horizontal axis 22 corresponds to time so that samples 18 forming one column, or being horizontally aligned, are samples belonging to one common time instant. Such set / column of temporally co-located samples 18 is illustrated in Fig. 5 at 24.
[0202] Each channel, thus, forms a digital time-varying signal ortime / amplitude or time-to-amplitude signal. The multi-channel digital signal m might have been obtained by at least one of Electrocardiography, Electroencephalography, Electromyography or seismic measurement. Differently speaking, the multi-channel digital signal might be a bio-physiological waveform data such as an electroencephalography (EEG) signal, an electrocardiogram (ECG), or an electromyography (EMG) signal, or seismic waveform data. However, each channel / signal might alternatively be another sort of waveform signal data such as scalar media data such as an audio signal and the signal 14 might be a multi-channel audio signal. Fig. 5 illustrates the option according to which signal 14 is not coded directly, i.e., in the original domain 26, but in a so-called “coded domain” 28 which might differ from the original domain 26 by one or more of 1) channel transformation, 2) channel permutation and 3) temporal mutual channel alignment. The channel transformation, if applied, transforms, per sample time instant, a set or column 24 of samples from domain 26 to domain 28. Thus, in domain 28, the sample pitch and the time axis is the same as in domain 26, but the meaning of the channels is different, i.e., the “source” channels of domain 26 become transformed channels in domain 28. Accordingly, the vertical axis in Fig. 5 for domain 28 is denoted as 32. Note that the channel transformation might leave the number of channels unchanged so that there is the same number of channels in domain 26 as well as domain 28, but different approaches are also possible. Generally, the channel transformation would aim at reducing redundancy and trying to condense the channels’ energy onto a fewer number of channels in domain 28. As said, the channel transformation is optional. Accordingly, in general terms, the channels in domain 28 are called “coded channels” in order to distinguish them from the “original” or “source” channels of digital signal 14 in domain 26. The permutation is also optional and may be used in combination with, or without, the channel transformation. If used in combination with the channel transformation, the permutation may be performed prior to and / or or subsequent to the channel transformation in order to permute / sort the source channels prior to transformation and the coded channels subsequent to the channel transformation. The channel transformation might be a DCT, DST, FFT or any other transformation. The temporal mutual alignment is also optional and might be seen as a constant temporal alignment between the source channels or the coded channels.
[0203] The module in encoder 10 performing the one or more of channel transformation, channel permutation and temporal mutual alignment is indicated in Fig. 5 as block 34. Side information 36 might be used in order to signal information on one or more of the following: 1) The channel transformation used, 2) information on the permutation(s) among the source channels and / or coded channels and 3) information on the mutual temporal alignment / delays between the source channels or coded channels wherein the temporal mutual alignment might be restricted to full sample precision. A corresponding block 38 in decoder 12 performs the reverse step, i.e., performs one or more of: 1) a channel retransformation, 2) a re-permutation of the source channels and / or coded channels and 3) a temporal re-alignment of the source channels or coded channels. Note, that if no channel transformation takes place, the coded channels are, in fact, equal to the source channels except for being temporally mutually aligned or being differently sorted due to permutation. Block 38 might be controlled by the before-mentioned side information 36. Thus, the “actual coding” relates to the coded channels in domain 28. In the coded domain 28, the coded channels are depicted in Fig. 5 as lines or rows of samples 40, each extending along time axis 22, the coded channels being depicted one on top of the other along coded channel axis 32 - potentially ordered according to a coded channel ID they have associated therewith - so as to result into an array of samples 40. Again, although Fig. 5 depicts the case that the number of source channels equals the number of coded channels, the number might be different. Further, if channel transformation is used, while there is no longer a clear association between source channels on the one hand and coded channels on the other hand, the temporal association remains: For each temporally co-located samples 24, there is a corresponding temporally co-located set 42 of samples 40 of the coded channels, wherein the set 42 in domain 28 is a column and might be a set of horizontally mutually offset samples in case of, and according to, the mutual temporal alignment, if applied. In case of Fig. 5, it has been assumed that no such temporal alignment took place so that both sets 42 and 24 are pure columns in the time / channel representation.
[0204] The actual coding is done in units of so-called temporal blocks 30. The term “block” or “temporal block” 30 is used so as to denote both a temporal portion of the multi-channel signal in domain 28, i.e., the set of coded channels, as well as a temporal portion of a certain coded channel. That is, for each temporal block 30, each coded channel has a temporal block such as block 540 depicted for some temporal block 30c and same are mutually co-located. The coding is done sequentially along these blocks 540, by following a coding / decoding order, which traverses the blocks 540 temporal block 30 by temporal block 30 with traversing temporally co-located blocks of the coded channels along a channel order corresponding to the order of the coded channels along axis 32. This coding / decoding order is illustrated in Fig.
[0205] 5 at 60. That is, in case of temporal block 540 being the block currently to be coded / decoded, the previously decoded / encoded temporal blocks include all preceding temporal blocks of all coded channels as well as the temporally co-located temporal blocks of coded channels preceding the coded channel 92 of temporal block 540 in channel order. These previously coded / decoded temporal blocks and their samples are illustrated in Fig. 5 by way of shading. In this regard, note that in Fig. 5, merely one temporal block 540 has been illustrated explicitly in order to reduce the complexity of Fig. 5. Thus, in the specification herein, reference sign 540 is sometimes used to indicate the currently encoded / decoded temporal block or to stand representatively for all temporal blocks. Further, as depicted in Fig. 5, the partitioning of signal 14 into temporal blocks 30 and 540, respectively, might be done in a manner so that these blocks 30 and 540, respectively, are non-overlapping. The actual coding in units of the temporal blocks 540 is performed predictively. That is, the encoder 10 comprises a block predictor 62 which predicts the samples of the currently coded temporal block 540, thereby yielding a prediction signal 64, and the prediction residual 66 formed by a subtraction between the actual sample values of temporal block 540 and the predicted samples of prediction signal 64 formed at a subtractor 68 is coded into the datastream 16 by residual coder 70. The residual coding in residual coder 70 may, or may not, involve a coding error by means of quantization. In any case, block predictor 62 uses the reconstructable version as being available by previously coded temporal blocks in order to obtain the prediction signal 64. This reconstructable version 72 might be derived at encoder 10 by means of a residual decoder 74 which reverses, potentially under coding loss, such as quantization, e.g. by means of dequantization, the residual signal 76 as coded into datastream 16, and an adder 78 which sums-up prediction signal 64 and the reconstructable residual signal 80 as obtained by residual decoder 74. To be more precise, let’s call the channel-individual temporal blocks 540 subblocks with temporally collocated subblocks of all channels forming a temporal block 30. Then, the prediction in module 62 or, to be more precise, the prediction at encoder and decoder, is performed in units of the subblocks 540, i.e. subblock wise. The encoder is free to choose different prediction modes for the subblocks within one block 30. As explained in more detail herein, within one block 30, one subblock 540 may be predicted based on one or more subblocks previously - according to the decoding order 60 - en / decoded within this block 30, while another subblock 540 within that block 30 might be coded / decoded based on the previously en / decoded subblock 540 of the same channel (but within the previous block 30). The transform residual en / decoding is then performed subblock wise by use of a onedimensional transform signaled in the data stream as described hereinbelow.
[0206] The decoder 12 decodes the coded channels from data stream 16 in a corresponding manner, i.e., in units of the temporal blocks 30 or in temporal blocks 540, respectively, and using predictive decoding. To this end, the decoder 12 comprises a residual decoder 82, an adder 84 and a block predictor 86 which correspond to, and are mutually connected in the same manner as, elements 74, 78 and 62 of encoder 10. That is, the residual decoder 82 derives from the residual signal 76 in data stream 16 the reconstructable residual signal 80 for a currently decoded temporal block 540 which is then subject to addition with prediction signal 64 derived by block predictor 86 for temporal block 540 on the basis of the reconstructed version 72 of previously decoded temporal blocks at adder 84. The output of adder 84, thus, yields the reconstructed version 72 of the currently decoded temporal block 540 and becomes part of the pool of already decoded samples of previously decoded temporal blocks when the temporal blocks of the coded channels are, in this manner, traversed along coding / decoding order 60 so as to reconstruct the coded channels in the coded domain 28. Note that the above description concentrated on the so-called sample prediction where samples of a current block 540 are predicted based on reconstructed samples of one or more previously decoded blocks, but coding inter dependencies, namely intra-channel and interchannel coding dependencies may be exploited not only in terms of sample prediction, but also in terms of other coding tools involving, for instance, parameter prediction and / or context derivation.
[0207] In order to enable a high degree of random access capability, some of the temporal blocks 30 may be coded in a random access manner meaning that the coded channels therein are coded independent from previous temporal blocks 30. Imagine, for instance, that temporal blocks 30b and 30e are random access temporal blocks. Then, none of the temporal channel blocks 540 in temporal block 30b as well as 30e would depend on any preceding temporal block 540 and no coding dependency would cross these temporal blocks 30b and 30e, that is no temporal block 540 within any of temporal block 30b-30d would be coded depending on any block 540 temporally preceding temporal block 30b, and no temporal block 540 within any of temporal block 30e and following would be coded depending on any block 540 temporally preceding temporal block 30e.
[0208] Thus, in other words, coding dependencies are restricted so as to not reach-out beyond the border of a random access temporal block 30b and 30e towards any preceding temporal block 30. Such restriction might also hold for intermediate temporal blocks 30c to 30d between random access temporal blocks 30b and 30e in that same may not depend on any temporal block preceding the leading one among the random access temporal blocks 30b and 30e, here block 30b. Accordingly, leading temporal borders of the random access temporal blocks 30b and 30e are indicated by bold lines in Fig. 5. In a variant, the restriction is not valid for all en / decoding stages. For instance, while the grouping might hold true for prediction, but the residual en / decoding dependencies might cross borders between channel groups. It might be the case, for instance, that for the entropy coding and decoding, all channels are coded jointly, i.e. using a single arithmetic coding engine, but that for the sake of prediction and reconstruction, the channels are grouped as described into independent groups such that, after entropy decoding, each such group can be reconstructed completely independently from each other group. This means that no prediction of sample values or any other information is supported between different channel groups.
[0209] Further, it might be that the coding of the coded channels also interrupts or restricts interchannel dependencies. For example, one or more of the coded channels might be coded as random access coded channels so that same do not use inter-channel dependencies, but merely intra-channel dependencies. The restriction of inter-channel coding dependencies might follow the channel order 32: that is, coding of these random access coded channels and the intermediate coded channels therebetween would be restricted so as to not reach-out beyond such a random access coded channel toward any coded channel preceding that random access coded channel in channel order along axis 32. Two such random access coded channels 88a and 88b and the resulting inter-channel dependency borders are illustrated in Fig. 5. Note that the restriction of inter-channel dependencies might be differently and is illustrated here merely as an example where the definition of, along channel order 32, interspersed random access channels 88a and 88b defines channel groups covering contiguous channels along the channel order 32. Other groups of channels might be defined, which do not necessarily follow the channel order 32, and inter-channel dependencies might be restricted not to render any channel of one group dependent on a channel of any other group, and within each group the inter-channel dependencies may also by restricted or each channel might by coded inter-channel dependent on any previously coded channel within its channel group.
[0210] The block predictor 62 and 86 of encoder 10 and decoder 12, respectively, operate synchronously, i.e., they generate the same prediction signal 64 based on the previously encoded / decoded samples of previously encoded / decoded temporal blocks 540. On encoder side 10, the prediction fora certain temporal block 540 may be accompanied or determined by one or more prediction parameters. Same might be determined on encoder side based on a rate / distortion optimization. These prediction parameters 90 are coded into data stream 16 and they are decoded from data stream 16 and used by block predictor 86 so as to perform the same prediction.
[0211] It might be that encoder 10 and decoder 12 support more than one prediction mode. For instance, encoder 10 and decoder 12 may support an intra prediction mode (which mode may also be called block-copy mode) according to which the currently encoded / decoded temporal block 540 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of the same coded channel to which the currently encoded / decoded temporal block 540 belongs, which is coded channel 92 in the example of Fig. 5. Additionally or alternatively, encoder 10 and decoder 12 may support an inter-prediction mode (which mode may also be called cross-channel prediction mode) according to which the currently encoded / decoded temporal block 540 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of one or more coded channels preceding - in coding order 32 - the coded channel 92 to which the currently encoded / decoded temporal block 540 belongs. Additionally or alternatively, there may be a mixed prediction mode according to which the prediction signal 64 is obtained by both, reconstructed / reconstructable sample values of previously encoded / decoded temporal blocks of coded channel 92 itself as well as reconstructed / reconstructable sample values of one or more coded channels preceding coded channel 92 in channel order along axis 32. Beyond this, there may be temporal blocks 540 which are coded without any prediction at encoder 10 and decoded without any prediction at decoder 12 such as the first temporal blocks 540 in the tiles 94 resulting from mutually separating the temporal blocks by means of the random access borders 96 on the one hand and the random access channel borders 98 on the other hand. This corresponds to the prediction signal 64 being set to zero and this may form an additional mode which could be called bypass mode. Additionally, or alternatively, there may be other modes such as ones deriving a DC predictor or linear function predictor for block 64 based on immediately preceding samples which immediately precede block 540. The prediction parameters 90 may, thus, contain for a currently encoded / decoded temporal block 540 a prediction mode flag or prediction mode indicator indicating the prediction mode to be used for this currently encoded / decoded temporal block 540 and, optionally, one or more parameters parameterizing the prediction mode to be used for this currently encoded / decoded temporal block 540. It might also be that the prediction parameters are themselves coded predictively from already reconstructed blocks 540. In this prediction process, the laid out random-access capabilities in channel- and temporal-direction are, as an example, always maintained, i.e. the mentioned prediction of prediction parameters may never be supported across such a random access segment.
[0212] As mentioned, the aforementioned coding dependencies ought not to cross any of the borders 96 and 98 not only result from the just-described sample prediction capabilities of block predictor 62 and 86, respectively, but may optionally also result from other mechanisms such as parameter prediction according to which parameters such as the aforementioned prediction parameters 90 for a certain temporal block 540 are predicted based on coding parameters conveyed in the data stream 16 for any previous temporal block, or context derivation for context-adaptive entropy coding / decoding any coding parameter such as the prediction parameters 90 or any other side information such as side information 76 and 36 for temporal block 540 based on any coding parameter conveyed in the data stream 16 for any preceding temporal block.
[0213] That is, summarizing, the encoder 10 encodes the multi-channel signal 14 by transferring it into the coded domain 28 and then coding the coded channels into data stream 16 in the just-described block-wise and predictive manner, wherein decoder 12 decodes the coded channels of coded domain 28 from data stream 16 and the corresponding block-wise and predictive manner with then gaining the multi-channel signal 14 in its original form 26 based on the coded channels in coded domain 28 by means of segment 38. As said, the channel transformation is optional and if not used, each sample 40 in the coded domain 28 really corresponds to one sample 18 in the original domain 26. If, further, the temporal mutual alignment is not used, each sample 40 exactly corresponds to a sample 18 in the original domain 26 at exactly the same time instant or, differently speaking, all temporally co-located samples 40 in coded domain 28 remain mutually temporally co-located in the original domain 26.
[0214] It should be noted that the temporal blocks 30 might, other than illustrated in Fig.5, vary in block length rather than being of a constant length as depicted in Fig. 5. For instance, encoder 10 may decide on the length of blocks 30 and signal the block length of blocks 30 (and the corresponding temporal blocks 540 of the coded channels) within data stream 16. Such signaling might be done on block level, such as for each temporal block 30 or, differently speaking for each temporally aligned bundle of blocks 540, so that the encoder may decide on the block size on the fly, or the block length might be signaled in the stream 16 on a larger scope such as for a sequence of blocks or even the whole stream 16.
[0215] As to the residual coder and residual decoder 70 and 82, they may use transform coding / decoding in order to convey the residual signal 76 in data stream 16. That is, the residual signal 80 may be conveyed in data stream 16 in transform or spectral domain by way of transform coefficients in residual signal 76. The transform domain might be a DCT, DST or an FFT. The transform may be non-overlapping, i.e. it may only transform residual signal 80 and its re-transform may only cover residual signal 76 within block 540, and / or may be nonwindowed, i.e. the residual signal might be transformed without any transform window used to temporally shape the residual signal 80 before the transform. The transform domain, i.e. the transformation leading from time domain to transform domain which is used by the encoder to transform the prediction residual signal 80 to be coded und the corresponding retransformation leading from transform domain to time domain which is used by the decoder to derive the prediction residual signal 80, or the transformation, might be selected from a set of available transforms including, for instance, one or more of 1) one or more DCTs, 2) one or more DSTs and 3) an identity transform according to which the prediction residual signal 80 is coded into the data stream 14 in time domain directly. The transform may be critically sampled in that the number of transform coefficients resulting from the samples of one block 540 may equal the number of samples of block 540. Again, the samples might be the residual samples or may be, in case of the bypass mode, the channel samples directly. The transform coefficients might be encoded by quantization, i.e. they may be quantized with the quantized coefficients then being coded in the datastream 16. Dequantization may occur at decoding. For quantization, either a scalar uniform reconstruction quantizer or a low complexity vector quantizer might be used. In order to determine the quantization indices, the encoder may perform some optimization algorithm such as a rate-distortion optimized scalar quantization, or a trellis quantization with the goal to approximately minimize an approximated Lagrangian rate-distortion cost. At the decoder, the reconstruction process that yiels (e.g. yields) the transform coefficients may be conducted by multiplying the coded quantization indices with a certain step-size and, in case of the use of a low-complexity vector quantizer, by additionally invoking a state-machine based on the parity of previously decoded quantization indices in order to reconstruct the current quantization index.
[0216] In order to control the quantization noise, the transform coefficients might be subject to noise shaping. Spectral noise shaping may be used to shape the quantization noise spectrally. This may be done by signaling in the data stream spectral-band scale factors, i.e. a scale factor per spectral band, which represent a transfer function of a spectral filter which approximates the spectral envelope of the signal within the current block 540 (or its prediction residual, respectively), or signaling filter coefficients defining a temporal filter having a filter transfer function which approximates the spectral envelope of the signal within the current block 540 (or its prediction residual, respectively). On encoder side, spectral noise shaping may be applied in spectral domain by multiplying an inverse of scale factors, either directly signaled in the data stream or derivable from the filter coefficients by filter-to-factor conversion, with the transform coefficients before quantization. That is, at encoder, the coefficients are shaped by the inverse of the spectral envelope. At decoder side, spectral shaping may be applied in spectral domain by multiplying scale factors, either directly signaled in the data stream or derived from the filter coefficients by filter-to-factor conversion, with the transform coefficients, with then (for example with then the retransformation being performed based on the spectrally shaped transform coefficients). That is, at decoder, the coefficients are shaped by the spectral envelope before applying retransformation. Additionally or alternatively, temporal noise shaping might be applied. To this end, TNS filter coefficients might be determined and signaled by the encoder. The TNS filter coefficients may represent a transfer function which approximates the temporal envelope of the current block 540 (or its residual signal). The encoder may apply TNS filtering using the filter coefficients by spectrally filtering the possibly spectrally shaped transform coefficients so as to filter them with a transfer function corresponding to an inverse of the temporal envelope. The TNS filter coefficients might be derived by linear prediction analysis of the possibly spectrally shaped transform coefficients so as to derive a linear prediction filter, then used as TNS filter, which minimizes a prediction residual when spectrally applied on the possibly spectrally shaped transform coefficients. At the encoder, the TNS filtered coefficients are then quantized and entropy coded. At decoder side, the inverse takes place: the possibly spectrally shaped transform coefficients are inversely TNS filtered before applying retransformation. Additionally or alternatively, noise filling might be used. The filling may be applied to zero-quantized portions of the spectrum and controlled by the encoder via corresponding noise filling parameters.
[0217] As to the encoding / decoding the block or sequence of quantized transform coefficients of a current block into / from the data stream 16, arithmetic coding, such as context-adaptive binary arithmetic coding, CABAC, may be used. The CABAC encoding / decoding may by performed frame wise. That is, in each channel, the sequence of blocks 540 may be partitioned into immediately consecutive blocks 540, which form frames. This partitioning may be equal among the channels so that, again, a frame denotes both a temporal portion within each channel individually, as well as a temporal portion of the multi-channel signal, i.e. a collection of temporally aligned frames. Within each frame, the sequence of blocks 540 are CABAC en / decoded with once initializing the contexts and resetting the internal CABAC state at the beginning and then updating the contexts’ probabilities during en / decoding the respective frame. That is, blocks 540 are CABAC decodable merely in units of frames. The context initialization might be done independent from previous frames, or depending on the contexts as manifesting itself at the end of, of during, the en / decoding a previous frame.
[0218] Some deblocking processing might be used to avoid blocking artifacts. If, alternatively, an overlapped transform is used, an overlap-add processing with re-transforms of immediately preceding / succeeding temporal blocks of the same coded channel might be used in order to completely reconstruct the current temporal block’s 540 residual signal 76.
[0219] Besides such transform-(residual)-coded blocks there might be temporal blocks 540 which, additionally or alternatively, are coded using, besides the block prediction by block predictor 62 / 86 - which could be called a primary prediction - a secondary sample-wise prediction of the residual samples in residual block 66 such as by predicting a current sample’s residual sample by means of already decoded values of preceding - in sample coding order - residual samples in block 66 or 80, with then correcting same by means of a secondary-prediction-residual sample decoded from the data stream 16. The secondary-prediction-residual samples for such a block may coded into the data stream en block in a transform domain or samplewise in time domain. Note that the afore-mentioned spectral shaping of the residual signal of a block 540 might be seen as a sample wise residual prediction, i.e. the case where filter coefficients are signaled for a block which define a temporal filter having a filter transfer function which approximates the spectral envelope of the residual signal within a current block 540. In sample wise residual prediction, the residual predictor on a current block 540 might either be chosen out of a fixed set of prediction modes, where an index to such a residual prediction mode is signaled in the bit-stream, or the residual prediction mode might be ‘signal adaptive’. In the latter case, prediction filter coefficients for the residual predictor are determined at the encoder by solving for example a linear equation, and are then quantized and transmitted to the decoder. At the decoder, the coefficients are inverse quantized and then the sample-wise prediction is conducted with these coefficients. The number of used coefficients may vary per block and might also be signaled in the bit-stream. Additionally, it might optionally (i.e. indicated by some information in the bit-stream) be supported to invoke collocated samples from a previous block for the sample wise residual prediction. Finally, the coefficients of the sample wise residual prediction might be coded predictively, i.e. be predicted from used coefficients of a previous block, where only the differences to the current coefficients are transmitted.
[0220] A final note shall be made with respect to the juxtaposition of frames, blocks 540, channels and channel groups and regarding decoding order. The description above already described the fact that the channels might be grouped into channel group with each channel group being coded independently from each other, meaning that the blocks 540 in a certain channel group are coded without dependencies from channels outside their channel group. The decoding order 60, thus, would traverse the channels channel-group individually, channel group by channel group. Within each channel group, the blocks 540 are traversed as described: all temporally aligned blocks 540 of all channels fist, then proceeding with the next blocks 540 and so forth. A frame may have a sequence of blocks of a channel group encoded thereinto along the mentioned decoding order order, such as n temporally consecutive blocks 540 for all channels of a channel group. IF the channel group had m channels, m*n block104 would, thus, be coded into the frame. As mentioned, there might be dependent frames, for which the CABAC contexts are adopted from the preceding frame of the same channel group, i.e. the one having encoded the immediately preceding block 540. For such dependent frames, not only CABAC contexts may be adopted from the preceding frame, but it may also be allowed to allow for prediction from the preceding frame to the dependent frame. Prediction, and possibly also any coding dependencies, towards channels outside the channel group and, within the channel group, towards frames temporally preceding the mostly recently previously en / decoded independent frame would be disallowed. Thus, each tile shown in Fig. 5 by bold lines may represent a sequence of an independent frame flowed by zero, one or more dependent frames.
[0221] As mentioned before, Fig. 5 only represents a possible “framework” into which the previously described embodiments and the embodiments described subsequently may be built into. Many modifications may be performed with respect to Fig. 5, and some of these modifications might be mentioned in the subsequent description with respect to certain ones of the subsequently described embodiments, but these modifications shall then be treated as being also applicable with respect to other ones of the subsequently described embodiments.
[0222] The description is now resumed with respect to the announced subsequently described embodiments. However, as a note, optionally, encoder 10 of Fig. 5 may correspond to the encoder shown in Fig. 3 or Fig. 4. That is encoder 10 may, for example, comprise the respective features, functionalities and details of encoder 300 or 400 both individually or taken in combination. In particular, respective encoders 300 or 400 implemented in the framework shown in Fig. 5 or Fig. 6 may comprise, e.g. additionally, or in an according manner, some or all of the above discussed functionality of encoder 10. The same applies accordingly to decoder 12 with respect to the decoders of Fig. 1 and 2.
[0223] Next, reference is made to Fig. 6. Fig. 6 shows a schematic view of a concept according to embodiments of the invention. Fig. 6 shows an example for time differential coefficient coding in a given channel c of the waveform signal. As an example, the previously discussed percept_mode=2.
[0224] Hence, as an example, the previously discussed mode information may comprise the high-level syntax element percept_mode having a value of 2, hence indicating that a perceptual mode is selected (e.g. TNS or NS mode).
[0225] Furthermore, a parameter mode signal, e.g. in the form of a delta_time_flag may be present. Fig. 6 shows two cases, namely delta_time_flagf=O, see 610, e.g. indicating that a non-predictive mode is selected and delta_time_flagf=1, see 620, e.g. indicating that a predictive mode is selected
[0226] On axis 630 time is shown, with processing steps for a previous frame or block f-1, and a current frame or block f.
[0227] Hence, for the case of delta_time_flagf=0, in step f-1 filter coefficients 640, e.g. for TNS or NS (as an example of order L=2), AM and BM may be present (with C and DMbeing zero or not considered). For the next step, signal coefficients 650, of order L=4 may be desired, in the form of Af, Bf, Cf and Df. Hence, in the non-predictive mode, said new coefficients 650 may be fully transmitted, so that a frame independent reconstruction of coefficients is achieved.
[0228] For the case of delta_time_flagf=1, in step f-1 filter coefficients the same TNS or NS filter coefficients 640 are considered to be present and for step f, the same filter coefficients 650 are considered to be desired. However, in this case not the full set of new desired coefficients is signaled, but only a residual information 660, namely information Af-AM, Bf-BM, Cf-0 and Df-0, in order to achieve a frame dependent reconstruction via addition, see 670, (of the residual information with the filter coefficients for the previous step f-1).
[0229] It is to be noted that all features, functionalities and details shown in Fig. 6 are optional and may hence be incorporated individually or in combination in any of the embodiments disclosed herein.
[0230] Embodiments according to the invention comprise a decoder for decoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) from a data stream using block-based transform decoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), wherein the decoder is configured to decode a spectrum (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum) for a current (e.g. a currently considered block, e.g. a currently encoded block) block of the DWS from the data stream, decode a mode signal (e.g. percept_mode) from the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the decoder is configured to subject the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) (e.g. to decode the current block by subjection the spectrum to TNS), in case of the selected mode being the TNS mode; and wherein the decoder is configured to subject the spectrum to NS (e.g. to decode the current block by subjection the spectrum to NS), in case of the selected mode being the NS mode.
[0231] Further embodiments according to the invention comprise an encoder for encoding a digital waveform signal, DWS, (e.g. a biophysical, e.g. a medical, e.g. a geophysical, e.g. a seismic, e.g. an acoustic, e.g. an audio signal) into a data stream using block-based transform encoding (wherein, for example, a block corresponds to a frame, or wherein, for example, a frame comprises a sequence of blocks (e.g. temporally adjacent blocks), for example, wherein a block corresponds to a window, e.g. a window of a frame, or wherein a block describes a channel or a portion of a channel within one frame), wherein the encoder is configured to encode a current block (e.g. a currently considered block, e.g. a currently encoded block) of the DWS into the data stream by deriving a spectrum of the current block (e.g. a spectral domain representation, e.g. a transform domain representation, e.g. a cepstrum), encode a mode signal (e.g. percept_mode) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode (e.g. via a value 1) and a noise synthesis, NS, mode (e.g. via a value 0); wherein the encoder is configured to, in encoding the current block, encode the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering (e.g. in a sense of a prediction filtering in a frequency direction, e.g. over frequency, e.g. across frequency e.g. using a frequency domain prediction filtering, e.g. in a sense of using a filtering in which a prediction value obtained on the basis of one or more previous (e.g. frequency) spectral bin values (e.g. previously processed frequency bin values) is subtracted from (or generally: combined with) a current (e.g. currently considered) spectral (e.g. frequency) bin value, to obtain a processed current spectral, e.g. frequency, bin value) (e.g. decoding the current block by subjection the spectrum to TNS), in case of the selected mode being the TNS mode; and subjecting the spectrum to NS (e.g. decoding the current block by subjection the spectrum to NS), in case of the selected mode being the NS mode.
[0232] It was recognized that a mutually exclusive application of TNS or NS allows reducing a number of parameter sets to be transmitted from encoder to decoder. It was recognized that TNS and NS may have substantially non-overlapping target applications, such as low-energy, nontransient frames, wherein NS may be more important (e.g. having a higher impact on a respective hearing impression), and higher-energy, often transient frames, wherein TNS may be more important (e.g. having a higher impact on a respective hearing impression). Accordingly, a mutually exclusive signalling and application of TNS and NS may be performed with only limited impact on hearing quality but with a significant impact on the signalling effort, hence a reduction of said effort. In addition, it was recognized that a common filtering structure, or in other words a same kind of filter parameters, can be used both for TNS and NS. Hence, it was recognized that respective parameters may be encoded and respectively decoded depending on one another. In other words, based on a set of transmitted TNS parameters, NS parameters may be obtained, e.g. predicted or vice versa. In particular, a filter parameter prediction may be performed mode agnostic, e.g. irrespective of whether NS or TNS is to be applied and irrespective of whether in a preceding decoding / encoding step (e.g. for a previous block or frame) TNS or NS parameters were decoded / encoded.
[0233] This may allow a very efficient manner of providing a parametrization for TNS and NS functionalities. Optionally, a high level syntax element, e.g. percept_mode may be used in order to signal whether any of TNS or NS parameters are present in a respective encoded representation, or correspondingly, whether any of TNS or NS is to be applied to a spectrum.
[0234] If such a high level syntax element (e.g. high level referring to the syntax element being transmitted once per entire DWS or at least once per channel) has a predetermined value, a respective encoded representation may comprise a mode signal, e.g. another syntax flag, e.g. a block_percept_mode, indicating in a mutually exclusive manner the TNS mode or the NS mode. Otherwise, the mode signal may not be present and hence a respective decoding / encoding may be skipped.
[0235] As another optional feature, a parameter mode information, e.g. another flag, such as delta_time_flag, may be encoded, enabling or disabling a respective differential prediction mode, e.g. as shown in Fig. 1. Hence, TNS parameters or NS parameters may be predicted irrespective of whether preceding parameters were TNS or NS parameters.
[0236] In general, mode signal and parameter mode information may be provided on a block-basis, on a frame-basis, on a channel basis, or on a per frame in channel basis, hence reducing a signalling effort depending on a granularity of the signalling.
[0237] Furthermore, it is to be noted, that according to embodiments, a NS-level may be transmitted (e.g. optionally even always be transmitted) in addition to NS filter parameters. In other words, in case of the selected mode being the NS mode, in a block / frame a noise level may be transmitted in addition.
[0238] Implementation alternatives Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0239] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0240] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0241] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine-readable carrier.
[0242] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0243] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0244] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.
[0245] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
[0246] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0247] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0248] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0249] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0250] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0251] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software.
[0252] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.
[0253] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein. References
[0254] [1] ISO / IEC International Standard 14496-3:2009, "Advanced Audio Coding," Geneva, 2009 or later.
[0255] [2] ISO / IEC International Standard 23003-3:2020, "Unified Speech & Audio Coding," Geneva, 2020.
Claims
Claims1. Decoder (12, 100, 200) for decoding a digital waveform signal, DWS, (251, 301, 401) from a data stream using block-based transform decoding,wherein the decoder is configured todecode a spectrum (111, 211) for a current block of the DWS from the data stream,decode a mode signal (112, 212) from the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode;wherein the decoder is configured tosubject the spectrum to TNS filtering,in case of the selected mode being the TNS mode; andwherein the decoder is configured tosubject the spectrum to NS,in case of the selected mode being the NS mode.
2. Decoder (12, 100, 200) of claim 1 ,wherein the decoder is configured to decode filter parameters (113, 281, 216, 261) from the data stream for the current block;wherein the decoder is configured touse the filter parameters for the TNS filtering,in case of the selected mode being the TNS mode; andwherein the decoder is configured tofill zero-quantized portions of the spectrum using noise, andshape the noise using the filter parameters,in case of the selected mode being the NS mode.
3. Decoder (12, 100, 200) of claim 2, configured todecode the filter parameters (113, 281, 261) bypredicting the filter parameters based on filter parameters (28T) signaled in the data stream for a previous block to obtain predicted filter parameters (272), decoding a prediction residual (214), andcorrecting the predicted filter parameters using the prediction residual.
4. Decoder (12, 100, 200) of claim 3, configured todecode the filter parameters (113, 281, 261) bypredicting the filter parameters based on filter parameters (28T) signaled in the data stream for the previous block to obtain the predicted filter parameters (272), decoding the prediction residual (214) andcorrecting the predicted filter parameters using the prediction residual, irrespective of whether the filter parameters for the current block and the filter parameters signaled in the data stream for the previous block are associated with a same selected mode.
5. Decoder (12, 100, 200) according to any of claims 3 to 4, configured todecode a parameter mode information (213) from the data stream, which indicates a selected parameter mode out of a predictive mode and a non-predictive mode;decode the filter parameters (113, 281, 261) bypredicting the filter parameters based on filter parameters (28T) signaled in the data stream for the previous block of the DWS to obtain the predicted filter parameters (272),decoding the prediction residual (214) and correcting the predicted filter parameters using the prediction residual,in case of the selected parameter mode being the predictive mode;decode the filter parameters (113, 216, 261) from the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS, in case of the selected parameter mode being the non-predictive mode.
6. Decoder (12, 100, 200) according to claim 5,wherein the parameter mode information (213) is a block-specific and / or channel specific information; andwherein the decoder is configured to selectively decode the filter parameters using the predictive mode or the non-predictive mode on a per-block basis and / or on a perchannel basis.
7. Decoder (12, 100, 200) according to any of claims 2 to 6, configured todecode the filter parameters (113, 281, 216, 261) as reflection coefficients and / or as lattice coefficients.
8. Decoder (12, 100, 200) according to any of claims 3 to 7, configured toperform the prediction of the filter parameters (281’) and / or the correction of the predicted filter parameters (272) in a spectral domain.
9. Decoder (12, 100, 200) according to any of claims 1 to 8, configured tosubject the spectrum (111, 211) to TNS filtering, in case of the selected mode being the TNS mode and subject the spectrum to NS, in case of the selected mode being the NS mode,using a same filter structure.
10. Decoder (12, 100, 200) according to any of claims 1 to 9,wherein the decoder is configured todecode noise synthesis parameters (216, 217) from the data stream for the current block;fill zero-quantized portions of the spectrum (111, 211) using noise, and shape the noise using the noise synthesis parameters,in case of the selected mode being the NS mode.
11. Decoder (12, 100, 200) according to any of claims 1 to 10,wherein the decoder is configured to selectively switch betweendecoding filter parameters (113, 281, 216, 261), andusing the filter parameters for the TNS filtering,andperforming time-domain gain control in order to perform TNS,in case of the selected mode being the TNS mode.
12. Decoder (12, 100, 200) according to any of claims 1 to 11,wherein the decoder is configured to decode a high level syntax element, indicating whether the selected mode is a perceptual mode, which is to be considered for every block of the DWS.
13. Decoder (12, 100, 200) according to any of claims 1 to 12,wherein the decoder is configured to decode a channel-specific high level syntax element, indicating whether the selected mode is a perceptual mode, which is to be considered for every block of a channel of the DWS.
14. Decoder (12, 100, 200) according to any of claims 1 to 13,wherein the decoder is configured to decode a high level syntax element, indicating whether the selected mode is any out of the TNS mode or the NS mode; andwherein the decoder is configured to skip decoding the mode signal, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode.
15. Decoder (12, 100, 200) according to any of claims 1 to 14,wherein the mode signal (121, 212) comprises a m-ary syntax element, wherein a first value indicates the TNS mode and wherein a second value indicates the NS mode.
16. Decoder (12, 100, 200) according to any of claims 1 to 15,wherein the decoder is configured to decode filter parameters (113, 281, 216, 261) from the data stream for the current block;wherein the decoder is configured todecode a noise level information,in case of the selected mode being the NS modewherein the decoder is configured touse the filter parameters for the TNS filtering,in case of the selected mode being the TNS mode; andwherein the decoder is configured tofill zero-quantized portions of the spectrum using noise and shape the noise using the filter parameters and using the noise level information, in case of the selected mode being the NS mode.
17. Encoder (10, 300, 400) for encoding a digital waveform signal, DWS, (301, 401) into a data stream using block-based transform encoding,wherein the encoder is configured toencode a current block of the DWS into the data stream by deriving a spectrum (311 , 411) of the current block,encode a mode signal (321, 421) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode;wherein the encoder is configured to, in encoding the current block, encode the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering, in case of the selected mode being the TNS mode; andsubjecting the spectrum to NS, in case of the selected mode being the NS mode.
18. Encoder (10, 300, 400) of claim 17,wherein the encoder is configured to encode filter parameters (341, 381, 441, 481, 461) into the data stream for the current block;wherein thethe filter parameters are to be used at the decoder for the TNS filtering, in case of the selected mode being the TNS mode; andthe filter parameters are to be used for shaping, by the decoder, noise using which zero-quantized portions of the spectrum are to be filled at the decoder, in case of the selected mode being the NS mode.
19. Encoder (10, 300, 400) of any of claims 17 or 18,wherein the encoder is configured tosubject the spectrum (311, 411) of the current block to TNS filter analysis to obtain TNS filter parameters (341, 441); andapply the TNS filter parameters to the spectrum to obtain a TNS filtered spectrum (351, 451);encode the TNS filtered spectrum into the data stream, andencode a filter information (341, 441) associated with the TNS filter parameters, in case of the selected mode being the TNS mode; andwherein the encoder is configured todetermine noise synthesis parameters (381, 481, 482) on the basis of the spectrum of the current block, for subjecting the spectrum to NS in a decoder, encode the spectrum, andencode a noise synthesis information (381, 481, 482) associated with the noise synthesis parameters,in case of the selected mode being the NS mode.
20. Encoder (10, 300, 400) of any of the claims 17 to 19,wherein the encoder is configured to selectively switch between the TNS mode and the NS mode as the selected mode;wherein the encoder is configured to subject the spectrum (311, 411) of the current block to TNS filtering using TNS filter parameters (341, 441), in order to obtain a TNS filtered spectrum (351, 451);wherein the encoder is configured to obtain noise synthesis parameters (381 , 481 , 482) on the basis of the spectrum of the current block, for subjecting the spectrum to NS in a decoder,wherein the encoder is configured toencode the TNS filtered spectrum, andencode a filter information (341, 441, 461, 472, 522) associated with the TNS filter parameters,in case of the selected mode being the TNS mode; andwherein the encoder is configured toencode the spectrum, andencode a noise synthesis information (381, 481, 482, 461, 472, 522) associated with the noise synthesis parameters,in case of the selected mode being the NS mode.
21. Encoder (10, 300, 400) according to claim 19 or 20,wherein the encoder is configured to obtain the noise synthesis parameters as noise synthesis filter parameters (481), andwherein the noise synthesis information comprises an information about the noise synthesis filter parameters.
22. Encoder (10, 300, 400) according to claim 21,wherein the encoder is configured topredict the TNS filter parameters based on filter parameters (521) signaled in the data stream for a previous block to obtain predicted TNS filter parameters (471),obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the TNS mode; andwherein the encoder is configured topredict the noise synthesis filter parameters based on filter parameters (521) signaled in the data stream for a previous block, to obtain predicted noise synthesis filter parameters (472),obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the NS mode.
23. Encoder (10, 300, 400) according to claim 22,wherein the encoder is configured to selectively switch between a predictive mode and a non-predictive mode as a selected parameter mode;wherein the encoder is configured topredict the TNS filter parameters based on filter parameters (521) signaled in the data stream for a previous block to obtain predicted TNS filter parameters (471),obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the TNS mode and in case of the selected parameter mode being the predictive mode; andwherein the encoder is configured topredict the noise synthesis filter parameters based on filter parameters (521) signaled in the data stream for a previous block, to obtain predicted noise synthesis filter parameters (471),obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the NS mode and in case of the selected parameter mode being the predictive mode;wherein the encoder is configured to encode the TNS filter parameters (441) in the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS, in case of the selected mode being the TNS mode and in case of the selected parameter mode being the non-predictive mode,wherein the encoder is configured to encode the noise synthesis filter parameters (481) in the data stream for the current block independent from filter parameters signaled in the data stream for the previous block of the DWS, in case of the selected mode being the NS mode and in case of the selected parameter mode being the non-predictive mode; andwherein the encoder is configured to encode a parameter mode information (421’) in the data stream, which indicates the selected parameter mode out of the predictive mode and the non-predictive mode.
24. Encoder (10, 300, 400) according to claim 23,wherein the parameter mode information (421’) is a block-specific and / or channel specific information; andwherein the encoder is configured to selectively encode the TNS filter parameters (441) and / or the noise synthesis filter parameters (481) using the predictive mode or the non- predictive mode on a per-block basis and / or on a per-channel basis.
25. Encoder (10, 300, 400) according to any of claims 22 to 24, configured toencode the TNS filter parameters (441) and / or the noise synthesis filter parameters (481) as reflection coefficients and / or as lattice coefficients or as prediction residuals thereof.
26. Encoder (10, 300, 400) according to any of claims 22 to 25, configured toperform the prediction of the TNS filter parameters (441) and / or of the noise synthesis filter parameters (481) and / or the obtaining of the prediction residual in a spectral domain.
27. Encoder (10, 300, 400) according to any of claims 21 to 26,wherein the encoder is configured topredict the TNS filter parameters (441) based on noise synthesis filter parameters (521) signaled in the data stream for a previous block, to obtain predicted TNS filter parameters (471),obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the TNS mode; andwherein the encoder is configured topredict the noise synthesis filter parameters (481) based on TNS filter parameters (522) signaled in the data stream for a previous block, to obtain predicted noise synthesis filter parameters,obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the NS mode.
28. Encoder (10, 300, 400) according to any of claim 21 to 27,wherein the encoder is configured topredict the TNS filter parameters (441) based on filter parameters (522) signaled in the data stream for a previous block, to obtain predicted TNS filter parameters (471), irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters,obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the TNS mode; andwherein the encoder is configured topredict the noise synthesis filter parameters (481) based on filter parameters signaled in the data stream for a previous block, to obtain predicted noise synthesis filter parameters (472), irrespective of whether the filter parameters signaled in the data stream for a previous block are noise synthesis filter parameters or TNS filter parameters,obtain a prediction residual (472), andencode the prediction residual;in case of the selected mode being the NS mode.
29. Encoder (10, 300, 400) according to any of claims 17 to 28,wherein the encoder is configured to selectively switch between encoding TNS filter parameters (441) and time-domain gain control parameters, in case of the selected mode being the TNS mode.
30. Encoder (10, 300, 400) according to any of claims 27 to 29,wherein the encoder is configured to encode a high level syntax element, indicating whether the selected mode is a perceptual mode, which is to be considered for every block of the DWS.
31. Encoder (10, 300, 400) according to any of claims 17 to 30,wherein the encoder is configured to encode a channel-specific high level syntax element, indicating whether the selected mode is a perceptual mode, which is to be considered for every block of a channel of the DWS.
32. Encoder (10, 300, 400) according to any of claims 17 to 31 ,wherein the encoder is configured to encode a high level syntax element, indicating whether the selected mode is any out of the TNS mode or the NS mode; andwherein the encoder is configured to skip encoding the mode signal, if the high level syntax element indicates that the selected mode is neither the TNS mode nor the NS mode.
33. Encoder (10, 300, 400) according to any of claims 17 to 32,wherein the mode signal comprises a m-ary syntax element, wherein a first value indicates that TNS mode and wherein a second value indicates the NS mode.
34. Encoder (10, 300, 400) according to any of claims 17 to 32,comprising a decoder according to any of claims D1 to D16, in order to predictively encode the DWS.
35. Encoder (10, 300, 400) according to any of claims 18 to 34,wherein the encoder is configured to obtain a noise level information,wherein the encoder is configured to obtain noise synthesis parameters as noise synthesis filter parameters (481), andwherein the noise synthesis information comprises an information about the noise synthesis filter parameters and the noise level information.
36. Method for decoding a digital waveform signal, DWS, (251, 301, 401) from a data stream using block-based transform decoding, the method comprising:decoding a spectrum (111 , 211) for a current block of the DWS from the data stream,decoding a mode signal (112, 212) from the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode;subjecting the spectrum to TNS filtering,in case of the selected mode being the TNS mode; andsubjecting the spectrum to NS,in case of the selected mode being the NS mode.
37. Method for encoding a digital waveform signal, DWS, (251, 301, 401) into a datastream using block-based transform encoding, the method comprising:encoding a current block of the DWS into the data stream by deriving a spectrum (311 , 411) of the current block,encoding a mode signal (321, 421) into the data stream, which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode;wherein the method comprises, in encoding the current block, encoding the current block so that the current block is to be decoded from the data stream by subjecting the spectrum to TNS filtering, in case of the selected mode being the TNS mode; andsubjecting the spectrum to NS, in case of the selected mode being the NS mode.
38. Computer program for performing the method according to claim 36 or 37, when the computer program runs on a computer.
39. Data stream comprising:an encoded representation of a spectrum (311, 411) for a current block of a DWS, andan encoded representation of a mode signal (321, 421), which mutually exclusively indicates a selected mode out of a temporal noise shaping, TNS, mode and a noise synthesis, NS, mode.
40. Data stream having encoded therein a digital wave form signal using the method according to claim 37.
Citation Information
Patent Citations
Audio encoder, audio decoder and related methods using two-channel processing within an intelligent gap filling framework
US20160210974A1
Audio processing method using complex number data, and apparatus for performing same
US20250104721A1
Audio processing method using complex number data, and apparatus for performing same
WO2023113490A1
Method and apparatus for spectrotemporally improved spectral gap filling in audio coding using a tilt
WO2023117144A1