Audio decoder, apparatus for determining a set of values defining characteristics of a filter, method for providing a decoded audio representation, method for determining a set of values defining characteristics of a filter, and computer program

By using machine learning-based filters to scale the spectral values ​​of the decoded audio signal in the audio decoder, the balance between bit rate, quality, and complexity in audio decoding technology is solved, thereby improving audio quality and computational efficiency.

CN114245919BActive Publication Date: 2025-12-05FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080035307.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-11
Filing Date
2020-04-09
Publication Date
2025-12-05
Estimated Expiration
2040-04-09

AI Technical Summary

Technical Problem

Existing audio decoding technologies struggle to find an effective balance between bit rate, audio quality, and complexity, resulting in poor audio decoding quality.

Method used

A machine learning-based filter is used to improve audio quality by scaling the spectral values ​​of the decoded audio signal. The filter is adjusted based on the spectral values ​​of the decoded audio representation, and the scaling value is determined using a neural network or machine learning structure, avoiding reliance on additional signaling information.

Benefits of technology

It improves audio decoding quality, reduces the impact of quantization noise, lowers computational complexity, achieves audio quality improvement in various situations, and has high computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114245919B_ABST
    Figure CN114245919B_ABST
Patent Text Reader

Abstract

An audio decoder for providing a decoded audio representation based on an encoded audio representation comprises a filter for providing an enhanced audio representation of the decoded audio representation. The filter is configured to obtain a plurality of scaling values associated with different frequency bands or frequency ranges based on spectral values of the decoded audio representation associated with the different frequency bands or frequency ranges, and the filter is configured to scale spectral values of the decoded audio signal representation or a pre-processed version thereof using the scaling values to obtain the enhanced audio representation. Also described is an apparatus for determining a set of values defining characteristics of a filter for providing an enhanced audio representation based on a decoded audio representation (122; 322).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments according to the application relate to an audio decoder.

[0002] Further embodiments according to the application relate to an apparatus for determining a set of values defining a characteristic of a filter.

[0003] Further embodiments according to the application relate to a method for providing a decoded audio representation.

[0004] Further embodiments according to the application relate to a method for determining a set of values defining a characteristic of a filter.

[0005] Further embodiments according to the application relate to a corresponding computer program.

[0006] Embodiments according to the application relate to a real-valued mask based postfilter for improving an encoded speech quantity.

[0007] Embodiments according to the application generally relate to a postfilter for enhancing decoded audio of an audio decoder, the postfilter being based on a decoded audio representation for determining a set of values defining a characteristic of a filter. BACKGROUND

[0008] Some conventional solutions will be introduced in the following.

[0009] In view of this situation, there is a need for a concept providing an improved trade-off between bit rate, audio quality and complexity when decoding an audio content. SUMMARY

[0010] Embodiments according to the application create an audio decoder (e.g. a speech decoder, or a general purpose audio decoder, or an audio decoder switching between a speech decoding mode (e.g. a linear prediction based decoding mode) and a general purpose audio decoding mode (e.g. a spectral domain representation based decoding mode using scaling factors for scaling decoded spectral values)) for providing a decoded audio representation based on an encoded audio representation.

[0011] The audio decoder comprises a filter (or "postfilter") for providing an enhanced audio representation (e.g. a speech representation) of the decoded audio representation (e.g. a speech representation) using the set of values defining the characteristic of the filter. )) provided by the decoder core of the audio decoder.

[0012] ​The filter (or postfilter) is configured to obtain a plurality of scaling values (e.g. mask values, e.g. M(k, n)) based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges (e.g. having a frequency band index or a frequency range index k), which scaling values may, for example, be real values and may, for example, be non-negative and may, for example, be limited to a predetermined range, and which are associated with different frequency bands or frequency ranges.

[0013] The filter (or postfilter) is configured to scale spectral values of the decoded audio signal representation (e.g. ) or a pre-processed version thereof using the scaling values (e.g. M(k, n)) to obtain an enhanced audio representation (e.g. ).

[0014] This embodiment is based on the idea that scaling of spectral values of the decoded audio signal representation can be used to effectively improve the audio quality, wherein the scaling values are derived based on spectral values of the decoded audio representation. It has been found that filtering affected by scaling of spectral values can effectively adapt to signal characteristics based on spectral values of the decoded audio representation and can improve the quality of the decoded audio representation. For example, based on spectral values of the decoded audio representation, filter settings (which can be defined by the scaling values) can be adjusted in a way that reduces the impact of quantization noise. For example, adjusting the scaling values based on spectral values of the decoded audio representation can use a machine learning structure or a neural network, which can provide the scaling values in a computationally efficient way.

[0015] In particular, it has been found that even though quantization noise is typically signal dependent, deriving the scaling values from spectral values of the decoded audio representation is still advantageous and can have good results. Thus, the concept can be applied and particularly good results can be obtained in this case.

[0016] In summary, the above described audio encoder allows to use a filter to improve the achievable audio quality, the characteristics of which are adjusted based on spectral values of the decoded audio representation, wherein the filtering operation can be performed in an efficient way, for example by scaling the spectral values using scaling values. Thus, the auditory impression can be improved, wherein it is not necessary to rely on any additional side information to control the adjustment of the filter. Instead, the adjustment of the filter can be based on the decoded spectral values of the current processing frame only, irrespective of the encoding scheme used to produce the encoded and decoded representations of the audio signal, and possible decoded spectral values of one or more previously decoded frames and / or one or more subsequently decoded frames.

[0017] In a preferred embodiment of the audio decoder, the filter is adapted to use a configurable processing structure (e.g. a "machine learning" structure like a neural network) in order to provide the scaling values, the configuration of which processing structure is based on a machine learning algorithm.

[0018] By using a configurable processing structure like a machine learning structure or a neural network, the characteristics of the filter can be easily adjusted based on coefficients defining the functionality of the configurable processing structure. Thus, the characteristics of the filter can typically be adjusted over a wide range depending on the spectral values of the decoded audio representation. Hence, an improved audio quality can be obtained in many different situations.

[0019] In a preferred embodiment of the audio decoder, the filter is configured to determine the scaling values based on the spectral values of the decoded audio representation only in the plurality of frequency bands or frequency ranges (e.g. without using any additional signaling information when deriving the scaling values from the spectral values).

[0020] Using such a concept, the audio quality can be improved independent of the presence of the side information.

[0021] Since a coherent and generic representation of the decoded audio signal (spectral values of the decoded audio representation) is used, the computational complexity and structural complexity can be kept at a reasonably low level, independent of the encoding technique used for obtaining the encoded representation and the decoded representation. In this case, complex and specific operations on specific side information values are avoided. Moreover, typically a generic processing structure (e.g. a neural network) can be used for deriving the scaling values based on the spectral values of the decoded audio representation, which uses a limited number of different computational functions (e.g. scaled summations, and evaluation of activation functions).

[0022] In a preferred embodiment of the audio decoder, the filter is configured to obtain the amplitude values of the enhanced audio representation according to (which can for example describe the absolute value or amplitude or norm)

[0023]

[0024] where M(k, n) is the scaling value, where k is a frequency index (e.g. specifying different frequency bands or frequency ranges), where n is a time index (which for example specifies different overlapping or non-overlapping frames), and where is the amplitude value of the spectral value of the decoded audio representation. The amplitude value may be the amplitude, absolute value, or any norm of the spectral value obtained by applying a time-frequency transform like a STFT (short-time Fourier transform), an FFT, or an MDCT to the decoded audio signal.

[0025] Alternatively, the filter can be configured to obtain the values of the enhanced audio representation according to

[0026]

[0027] wherein M(k, n) is a scaling value, wherein k is a frequency index (e.g. specifying different frequency bins or frequency ranges), wherein n is a time index (e.g. specifying different overlapping or non-overlapping frames), and wherein is a spectral value of the decoded audio representation.

[0028] It has been found that such a simple derivation of an enhanced amplitude value of the audio representation or of an (typically complex-valued) value of the audio representation can be performed with good efficiency and still results in a significant improvement of the audio quality.

[0029] In a preferred embodiment of the audio decoder, the filter is configured to obtain the scaling value such that the scaling value results in a scaling (or in some cases an amplification) of one or more spectral values of the decoded audio signal representation or of one or more pre-processed spectral values based on the spectral values of the decoded audio signal representation.

[0030] By performing such a scaling, an amplification or attenuation of at least one spectral value can preferably but not necessarily be caused (and typically also an attenuation of at least one spectral value can be caused) in an efficient manner. For example, by allowing both amplification and attenuation by scaling, artifacts that can be caused by a limited precision of the digital representation can in some cases also be reduced. Furthermore, by avoiding a restriction of the scaling value to values smaller than one, an additional degree of freedom for the adjustment of the scaling value is optionally included. Thus, a good improvement of the audio quality can be achieved.

[0031] In a preferred embodiment of the audio decoder, the filter comprises a neural network or a machine learning structure configured to provide the scaling value based on a plurality of spectral values describing the decoded audio representation (e.g. describing an amplitude of a transformed representation of the decoded audio representation), wherein the spectral values are associated with different frequency bins or frequency ranges.

[0032] It has been found that the use of a neural network or machine learning structure in such a filter leads to a relatively high efficiency. It has also been found that the neural network or machine learning structure can easily handle the decoding audio representation's spectral values of the input quantity in case the number of spectral values input to the neural network or machine learning structure is relatively large. It has been found that the neural network or machine learning structure can well handle such a large input signal or input quantity and that it can also provide a large number of different scaling values as output quantity. In other words, it has been found that the neural network or machine learning structure is very well suited to derive a relatively large number of scaling values based on a relatively large number of spectral values without requiring excessive computational resources. Thus, the scaling values can be adjusted in a very precise manner to the decoding audio representation's spectral values without an excessive computational load, wherein the details of the decoding audio representation's spectrum can be taken into account when adjusting the filter characteristics. Furthermore, it has been found that the coefficients of the neural network or machine learning structure providing the scaling values can be determined with reasonable effort and that the neural network or machine learning structure provides enough degrees of freedom to achieve a precise determination of the scaling values.

[0033] In a preferred embodiment of the audio decoder, the input signal of the neural network or machine learning structure represents the logarithmic amplitudes, amplitudes or norms of the decoding audio representation's spectral values, wherein the spectral values are associated with different frequency bins or frequency ranges.

[0034] It has been found to be advantageous to provide the logarithmic amplitudes of the spectral values, the amplitudes of the spectral values or the norms of the spectral values as input signal of the neural network or machine learning structure. It has been found that the sign or phase of the spectral values is secondary for the adjustment of the filter, i.e. for the determination of the scaling values. In particular, it has been found that it is particularly advantageous to logarithmize the amplitudes of the decoding audio representation's spectral values since the dynamic range can be reduced. It has been found that the neural network or machine learning structure can generally better handle the logarithmic amplitudes of the spectral values when compared to the spectral values themselves since the spectral values generally have a large dynamic range. By using logarithmic values, a simplified numerical representation can also be used in the (artificial) neural network or machine learning structure since the use of floating point numbers is generally not required. In contrast, the neural network or machine learning structure can be designed using fixed point numerical representations which significantly reduces the implementation effort.

[0035] In a preferred embodiment of the audio decoder, the output signal of the neural network or machine learning structure represents the scaling values, e.g. the mask values.

[0036] By providing the scaling values as output signal (or output quantity) of the neural network or machine learning structure, the implementation effort can be kept at a reasonably low level. For example, a neural network or machine learning structure providing a relatively large number of scaling values is easy to implement. For example, a homogeneous structure can be used which reduces the implementation effort.

[0037] In a preferred embodiment of the audio decoder, the neural network or the machine learning structure is trained to limit, reduce or minimize a deviation (e.g. a mean squared error; e.g. MSE MA ) between the plurality of target scaling values (e.g. IRM(k, n)) and the plurality of scaling values (e.g. M(k, n)) obtained using the neural network or using the machine learning structure.

[0038] By training the neural network or the machine learning structure in this way, it can be achieved that the enhanced audio representation (obtained by scaling spectral values (or pre-processed versions thereof) of the decoded audio signal representation using the scaling values) provides a good auditory impression. For example, the target scaling values can be easily determined, e.g. based on knowledge of the encoder-side lossy processing. Thus, it can be determined with little effort which scaling values best bring the spectral values of the decoded audio representation close to an ideal enhanced audio representation (e.g. which can be equal to the input audio representation of the audio encoder). In other words, by training the neural network or the machine learning structure to limit, reduce or minimize a deviation between the plurality of target scaling values and the plurality of scaling values obtained using the neural network or using the machine learning structure, e.g. for a plurality of different audio contents or different types of audio contents, it can be achieved that the neural network or the machine learning structure provides appropriate scaling values even for different audio contents or different types of audio contents. Moreover, by using the deviation between the target scaling values and the scaling values obtained using the neural network or using the machine learning structure as an optimization quantity, the complexity of the training process can be kept small and numerical problems can be avoided.

[0039] In a preferred embodiment of the audio decoder, the neural network or the machine learning structure is trained to limit, reduce or minimize a deviation (e.g. a mean squared error; e.g. MSE SA ) between the target amplitude spectrum, the target amplitude spectrum, the target absolute spectrum or the target norm spectrum (e.g. |X(k, n)|, e.g. the original spectrum of the training audio signal) and the (enhanced) amplitude spectrum, the amplitude spectrum, the absolute spectrum or the norm spectrum obtained using a scaling (e.g. a frequency-dependent scaling) of a processed (e.g. decoded, e.g. quantized, encoded and decoded) spectrum, the scaling of the processed spectrum (e.g. based on the target amplitude spectrum and / or based on the training audio signal) using the scaling values provided by the neural network or by the machine learning structure (wherein the input signal of the neural network is e.g. based on the decoded spectrum).

[0040] By using this training method, it is generally possible to ensure a good quality of the enhanced audio representation. In particular, it has been found that the neural network or machine learning structure also provides appropriate scaling factors if the decoded audio representation represents different audio content compared to the audio content used for training. Furthermore, it has been found that the enhanced audio representation is considered to have a good quality if the magnitude spectrum or amplitude spectrum or absolute spectrum or norm spectrum has a sufficiently good consistency with the desired (target) magnitude spectrum or (target) amplitude spectrum or (target) absolute spectrum or (target) norm spectrum.

[0041] In a preferred embodiment of the audio decoder, the neural network or machine learning structure is trained such that the scaling of one or more spectral values of a spectral decomposition of the decoded audio signal representation, or of one or more pre-processed spectral values based on the spectral values of the spectral decomposition of the decoded audio signal representation, is limited to a range between 0 and a predetermined maximum value.

[0042] It has been found that the limitation of the scaling (or scaling value) helps to avoid an excessive amplification of spectral values. It has been found that a very large amplification (or scaling) of one or more spectral values can lead to audible artifacts. Furthermore, it has been found that an excessive scaling value can be reached during training, for example, if a spectral value of the decoded audio representation is very small, or even equal to zero. Thus, the quality of the enhanced audio representation can be improved by using such a limitation method.

[0043] In a preferred embodiment of the audio decoder, the maximum value is larger than 1 (and can for example be 2, 5 or 10).

[0044] It has been found that such a limitation of the scaling (or scaling value) leads to particularly good results. For example, by allowing an amplification (for example, by allowing a scaling or scaling value larger than 1), artifacts caused by "spectral holes" can also be partially compensated. At the same time, excessive noise can be limited by attenuation (for example, using a scaling or scaling value smaller than 1). Thus, a very flexible signal improvement can be obtained by scaling.

[0045] In a preferred embodiment of the audio decoder, the neural network or machine learning structure is trained such that the scaling (or scaling value) of one or more spectral values of a spectral decomposition of the decoded audio signal representation, or of one or more pre-processed spectral values based on the spectral values of the spectral decomposition of the decoded audio signal representation, is limited to 2, or limited to 5, or limited to 10, or limited to a predetermined value larger than 1.

[0046] By using such a method, artifacts can be kept within a reasonably small range, while allowing an amplification (for example, this can help to avoid "spectral holes"). Thus, a good auditory impression can be obtained.

[0047] In a preferred embodiment of the audio decoder, the neural network or machine learning structure is trained such that the scaling values are limited to 2, or limited to 5, or limited to 10, or limited to a predetermined value larger than 1.

[0048] By limiting the scaling values to such a range, a particularly good quality of the enhanced audio representation can be achieved.

[0049] In a preferred embodiment of the audio decoder, the number of input features of the neural network or machine learning structure (e.g. 516 or 903) is at least twice as large as the number of output values of the neural network or machine learning structure (e.g. 129).

[0050] It has been found that using a relatively large number of input features (larger than the number of output values (or output signals) of the neural network or machine learning structure) of the neural network or machine learning structure leads to particularly reliable scaling values. In particular, by selecting a relatively large number of input features of the neural network, information from previous frames and / or from subsequent frames can be taken into account, wherein it has been found that taking such additional input features into account generally improves the quality of the scaling values, and thus the quality of the enhanced audio representation.

[0051] In a preferred embodiment of the audio decoder, the filter is configured to normalize the input features (e.g. represented by the input signal) of the neural network or machine learning structure (e.g. amplitudes of spectral values obtained using a short-time Fourier transform) to a predetermined mean value (e.g. a mean value of zero) and / or a predetermined variance (e.g. a variance of unity) or standard deviation.

[0052] It has been found that normalizing the input features of the neural network or machine learning structure makes the provision of the scaling values independent of the volume or loudness or intensity of the decoded audio representation. Thus, the neural network or machine learning structure can “focus” on structural features of the spectrum of the decoded audio representation and is not (or not significantly) influenced by volume variations. Furthermore, by performing such a normalization, it can be avoided that nodes of the neural network become overly saturated. Furthermore, the dynamic range is reduced, which helps to keep the numerical representations used in the neural network or machine learning structure valid.

[0053] In a preferred embodiment of the audio decoder, the neural network comprises an input layer, one or more hidden layers, and an output layer.

[0054] Such a structure of the neural network has proven advantageous for the present application.

[0055] In a preferred embodiment of the audio decoder, the one or more hidden layers use a rectified linear unit as activation function.

[0056] It has been found that using a rectified linear unit as activation function allows to provide the scaling vector based on the spectral values of the decoded audio representation with good reliability.

[0057] In a preferred embodiment of the audio decoder, the output layer uses a (unbound) rectified linear unit or a bound rectified linear unit or a sigmoid function (e.g. a scaled sigmoid function) as activation function.

[0058] By using a rectified linear unit or a bound rectified linear unit or a sigmoid function as activation function in the output layer, the scaling values can be obtained in a reliable way. Specifically, as mentioned above, using a bound rectified linear unit or a sigmoid function allows to limit the scaling values to a desired range. Thus, the scaling values can be obtained in an efficient and reliable way.

[0059] In a preferred embodiment of the audio decoder, the filter is configured to obtain short-time Fourier transform coefficients (e.g. MDCT coefficients) representing spectral values of the decoded audio representation associated with different frequency bins or frequency ranges. ).

[0060] It has been found that the short-time Fourier transform coefficients constitute a particularly meaningful representation of the decoded audio representation. For example, it has been recognized that in certain situations (even though the audio decoder can use MDCT coefficients to reconstruct the decoded spectral representation) a neural network or machine learning structure can work better using the short-time Fourier transform coefficients than using the MDCT coefficients.

[0061] In a preferred embodiment of the audio decoder, the filter is configured to derive a logarithmic amplitude, an amplitude, an absolute value or a norm value (e.g. based on the short-time Fourier transform coefficients) and to determine the scaling values based on the logarithmic amplitude, the amplitude, the absolute value or the norm value.

[0062] It has been found that deriving the scaling values based on non-negative values (e.g. logarithmic amplitude values, amplitude values, absolute values or norm values) is efficient since the consideration of the phase would significantly increase the computational demand without bringing any substantial improvement of the scaling values. Thus, removing the sign and typically also the phase of the spectral values (e.g. obtained by the short-time Fourier transform) brings a good trade-off between complexity and audio quality.

[0063] In a preferred embodiment of the audio decoder, the filter is configured to determine a plurality of scaling values associated with a current frame based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges of the current frame (e.g. of the decoded audio representation or of the short-time Fourier transform) and based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges of one or more frames preceding the current frame (e.g. past context frames).

[0064] However, it has been found that taking into account spectral values of one or more frames preceding the current frame helps to improve the scaling vectors. This is due to the fact that many types of audio content comprise a temporal correlation between subsequent frames. Thus, the neural network or machine learning structure can take into account the temporal evolution of the spectral values, e.g. when determining the scaling values. For example, the neural network or machine learning structure can adjust the scaling values to avoid (or counteract) an excessive variation of the scaled spectral values over time (e.g. in the enhanced audio representation).

[0065] In a preferred embodiment of the audio decoder, the filter is configured to determine the plurality of scaling values associated with the current frame based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges of one or more frames (e.g. future context frames) succeeding the current frame (e.g. a current frame of the decoded audio representation, or a current frame of the short-time Fourier transform).

[0066] By taking into account spectral values of the decoded audio representation of one or more frames succeeding the current frame, also the correlation between subsequent frames can be exploited, and generally the quality of the scaling values can be improved.

[0067] An embodiment according to the present application creates an apparatus for determining a set of values defining characteristics of a filter (e.g. a neural network-based filter, or a filter based on another machine learning structure) for providing an enhanced audio representation (e.g. an up-mix signal) based on a decoded audio representation (e.g. can be provided by an audio decoding) (e.g. a down-mix signal). ).

[0068] The apparatus is configured to obtain spectral values (e.g. amplitudes or phases or MDCT coefficients represented by amplitude values (e.g. magnitude values) associated with different frequency bands or frequency ranges of the decoded audio representation. ).

[0069] The apparatus is configured to determine the set of values defining characteristics of the filter such that scaling values provided by the filter based on the spectral values of the decoded audio representation associated with different frequency bands or frequency ranges approximate target scaling values (which can be computed based on a comparison of a desired enhanced audio representation and the decoded audio representation).

[0070] Alternatively, the apparatus is configured to determine the set of values defining characteristics of the filter such that a spectrum obtained by the filter based on the spectral values of the decoded audio representation associated with different frequency bands or frequency ranges and using scaling values obtained based on the decoded audio representation approximates a target spectrum (which can correspond to a desired enhanced audio representation, and which can be equal to an input signal of an audio encoder in a processing chain comprising the audio encoder and the audio decoder comprising the filter).

[0071] Using such an apparatus, a set of values defining characteristics of a filter (used in the above-mentioned audio decoder) can be obtained with a moderate effort. In particular, the set of values defining characteristics of the filter (which can be coefficients of a neural network or coefficients of another machine learning structure) can be determined such that the filter produces good audio quality using the scaling values and results in an improvement of the enhanced audio representation relative to the decoded audio representation. For example, the set of values defining characteristics of the filter can be determined based on a plurality of training audio contents or reference audio contents, wherein the target scaling values or the target spectrum can be derived from the reference audio contents. However, it has been found that the set of values defining characteristics of the filter is also very well suited for audio contents that are different from the reference audio contents, provided that the reference audio contents at least to some extent represent audio contents to be decoded by the above-mentioned audio decoder. Moreover, it has been found that using the scaling values provided by the filter or using the spectrum obtained by the filter as an optimization quantity results in a reliable set of values defining characteristics of the filter.

[0072] In a preferred embodiment of the apparatus, the apparatus is configured to train a machine learning structure (e.g. a neural network) (as part of the filter and to provide scaling values for scaling the amplitude values of the decoded audio signal, or the spectral values of the decoded audio signal) to reduce or minimize a deviation (e.g. a mean squared error; e.g. MSE MA ) between a plurality of target scaling values (e.g. IRM(k, n)) and a plurality of scaling values (e.g. M(k, n)) obtained using the neural network based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges.

[0073] By training the machine learning structure using the target scaling values, which can be derived, for example, based on an original audio content that is encoded and decoded in a processing chain comprising the audio decoder (which derives the decoded audio representation), the machine learning structure can be designed (or configured) to at least partially compensate for signal impairments in the processing chain. For example, the target scaling values can be determined such that the target scaling values scale the decoded audio representation in a way that the decoded audio representation approximates the (original) audio representation input into the processing chain (e.g. into the audio encoder). Thus, the scaling values provided by the machine learning structure can have a high degree of reliability and can be adapted to improve the reconstruction of audio contents that undergo the processing chain.

[0074] In a preferred embodiment, the apparatus is configured to train the machine learning structure (e.g. a neural network) to reduce or minimize a deviation (e.g. a mean squared error; e.g. MSESA ), the scaling of the processed spectrum (e.g. based on the target amplitude spectrum and / or based on the training audio signal) uses scaling values provided by a machine learning structure (e.g. a neural network). For example, the input signal of the machine learning structure or neural network is based on the decoded spectrum.

[0075] It has been found that such training of the machine learning structure also results in scaling values that allow to compensate for signal impairments in a signal processing chain (which can include audio encoding and audio decoding). For example, the target spectrum can be the spectrum of a reference audio content or training audio content that is input in a processing chain including a (lossy) audio encoder and an audio decoder (providing a decoded audio representation). Thus, the machine learning structure can be trained such that the scaling values scale the decoded audio representation to approximate the reference audio content input to the audio encoder. Thus, the machine learning structure can be trained to provide scaling values that help to overcome impairments in the (lossy) processing chain.

[0076] In a preferred embodiment, the apparatus is configured to train the machine learning structure (e.g. a neural network) such that the scaling for the spectral values of the decoded audio signal representation, or for one or more pre-processed spectral values based on the spectral values of the decoded audio signal representation, lies in a range between 0 and 2, or in a range between 0 and 5, or in a range between 0 and 10, or in a range between 0 and a maximum value (e.g. which can be larger than 1).

[0077] By limiting the scaling to be within a predetermined range (e.g. between zero and a predetermined value, which can typically be larger than 1), artifacts that can be caused by e.g. too large scaling values can be avoided. Moreover, it should be noted that the limitation of the scaling values (which can be provided as output signal of the neural network or machine learning structure) allows to relatively simply implement the output stage (e.g. output node) of the neural network or machine learning structure.

[0078] In a preferred embodiment of the apparatus, the apparatus is configured to train the machine learning structure (e.g. a neural network) such that the amplitude scaling for the spectral values of the decoded audio signal representation, or for one or more pre-processed spectral values based on the spectral values of the decoded audio signal representation, is limited to lie in a range between 0 and a predetermined maximum value.

[0079] By limiting the amplitude scaling (or scaling values) to lie in a range between zero and a predetermined maximum value, an impairment switch caused by too strong amplitude scaling is avoided.

[0080] In a preferred embodiment of the audio decoder, the maximum value is larger than 1 (and can e.g. be 2, 5 or 10).

[0081] By allowing the maximum value of the amplitude scaling to be larger than 1, both attenuation and amplification can be achieved by using the scaling of the scaling values. It has been shown that such a concept is particularly flexible and leads to particularly good auditory impressions.

[0082] One embodiment of the invention creates a method for providing an enhanced audio representation (e.g. ) based on a decoded audio representation (e.g.

[0083] ). The method comprises providing an enhanced audio representation (e.g. ) of a decoded audio representation (e.g. ), wherein the input audio representation used by the filter providing the enhanced audio representation can be provided, for example, by a decoder core of an audio decoder.

[0084] The method comprises obtaining a plurality of scaling values (e.g. mask values, e.g. M(k, n)) based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges (e.g. having a frequency band index or a frequency range index k), the scaling values can be, for example, real values, and can be, for example, non-negative, and can be, for example, limited to a predetermined range, and are associated with different frequency bands or frequency ranges (e.g. having a frequency band index or a frequency range index k).

[0085] The method comprises scaling spectral values (e.g. ) of the decoded audio signal representation or a pre-processed version thereof using the scaling values (e.g. M(k, n)) to obtain the enhanced audio representation (e.g. ).

[0086] The method is based on the same considerations as the above-described apparatus. Furthermore, it should be noted that the method can also be supplemented by any of the features, functions and details described herein with respect to the apparatus. Furthermore, it should be noted that the method can be supplemented individually and in combination by any of these features, functions and details.

[0087] An embodiment creates a method for determining a set of values defining a characteristic of a filter (e.g. a neural network-based filter, or a filter based on another machine learning structure) for providing an enhanced audio representation (e.g. ) based on a decoded audio representation (e.g. can be provided by an audio decoding).

[0088] The method comprises obtaining spectral values (amplitudes or phases or MDCT coefficients represented by amplitude values (e.g. ) associated with different frequency bands or frequency ranges of the decoded audio representation.

[0089] The method comprises determining a set of values defining characteristics of a filter such that a spectrum provided by the filter based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges and using scaling values obtained based on the decoded audio representation is close to a target spectrum (which can correspond to the desired enhanced audio representation and can be equal to an input signal of an audio encoder in a processing chain comprising the audio encoder and the audio decoder comprising the filter).

[0090] Alternatively, the method comprises determining a set of values defining characteristics of a filter such that a spectrum obtained by the filter based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges and using scaling values obtained based on the decoded audio representation is close to a target spectrum (which can correspond to the desired enhanced audio representation and can be equal to an input signal of an audio encoder in a processing chain comprising the audio encoder and the audio decoder comprising the filter).

[0091] The method is based on the same considerations as the above-described apparatus. However, it should be noted that the method can also be supplemented by any of the features, functions and details described herein with respect to the apparatus. Moreover, the method can be supplemented by features, functions and details individually and in combination.

[0092] A computer program according to an embodiment of the application creates, when the computer program is run on a computer, a computer program for performing the method described herein. BRIEF DESCRIPTION OF DRAWINGS

[0093] Embodiments according to the application will be described below with reference to the accompanying drawings, in which:

[0094] Figure 1 A schematic block diagram of an audio decoder according to an embodiment of the application is shown;

[0095] Figure 2 A schematic block diagram of an apparatus for determining a set of values defining characteristics of a filter according to an embodiment of the application is shown.

[0096] Fig. 3 shows a schematic block diagram of an audio decoder according to an embodiment of the application.

[0097] Fig. 4 shows a schematic block diagram of an apparatus for determining a set of values defining characteristics of a filter according to an embodiment of the application.

[0098] Fig. 5 shows a schematic block diagram of an apparatus for determining a set of values defining characteristics of a filter according to an embodiment of the application.

[0099] Table 1 shows a representation of the percentage of mask values lying within the interval (0, 1) for different signal-to-noise ratios (SNRs);

[0100] Table 2 shows a representation of the percentage of mask values in different threshold regions measured at the three lowest bitrates of AMR-WB;

[0101] Figure 6 a schematic representation of a fully connected neural network (FCNN) mapping log-magnitudes to real-valued masks is shown;

[0102] Figure 7 a graphical representation of average PESQ and POLQA scores evaluating oracle experiments with different mask binding values at 6.65 kbps is shown;

[0103] Figure 8 a graphical representation of average PESQ and POLQA scores evaluating the performance of the proposed method and the EVS post-processor is shown;

[0104] Figure 9 a flowchart of a method according to an embodiment of the application is shown; and

[0105] Figure 10 a flowchart of a method according to an embodiment of the application is shown. DETAILED DESCRIPTION

[0106] 1) An audio decoder according to Figure 1 the preamble of claim 1

[0107] Figure 1 a schematic block diagram of an audio decoder 100 according to an embodiment of the application is shown. The audio decoder 100 is configured to receive an encoded audio representation 110 and to provide an enhanced audio representation 112 based on the encoded audio representation 110, which can be an enhanced form of a decoded audio representation.

[0108] The audio decoder 100 optionally comprises a decoder core 120, which can receive the encoded audio representation 110 and provide a decoded audio representation 122 based on the encoded audio representation 110. The audio decoder further comprises a filter 130, which is configured to provide the enhanced audio representation 112 based on the decoded audio representation 122. The filter 130, which can be regarded as a post-filter, is configured to obtain a plurality of scaling values 136, which are also associated with different frequency bands or frequency ranges, based on spectral values 132 of the decoded audio representation, which are associated with different frequency bands or frequency ranges. For example, the filter 130 can comprise a scaling value determination or scaling value determinator 134, which receives the spectral values 132 of the decoded audio representation and provides the scaling values 136. The filter 130 is further configured to use the scaling values 136 to scale spectral values of the decoded audio signal representation or a pre-processed version thereof to obtain the enhanced audio representation 112.

[0109] It should be noted that the spectral values of the decoded audio representation used for obtaining the scaling values can be identical to the spectral values that are actually scaled (e.g., by the scaling or scaler 138) or can be different from the spectral values that are actually scaled. For example, a first subset of the spectral values of the decoded audio representation can be used for the determination of the scaling values and a second subset of the spectral values of the spectrum or the magnitude spectrum or the absolute spectrum or the norm spectrum can be actually scaled. The first subset and the second subset can be equal, or can partially overlap, or can even be completely different (without any common spectral values).

[0110] With respect to the functionality of the audio decoder 100, it can be said that the audio decoder 100 provides a decoded audio representation 122 based on an encoded audio representation. Since the encoding (i.e., the provision of the encoded audio representation) is typically lossy, the decoded audio representation 122 provided, for example, by the decoder core can comprise some deterioration compared to the original audio content that can have been fed into an audio encoder that provided the encoded audio representation 110. It should be noted that the decoded audio representation 122 provided, for example, by the decoder core can be in any form and can be provided, for example, by the decoder core in the form of a time domain representation or in the form of a spectral domain representation. The spectral domain representation can comprise, for example, (discrete) Fourier transform coefficients or (discrete) MDCT coefficients or the like.

[0111] The filter 130 can obtain (or receive), for example, spectral values representing the decoded audio representation. However, the spectral values used by the filter 130 can be, for example, of a different type than the spectral values provided by the decoder core. For example, the filter 130 can use Fourier coefficients as spectral values, whereas the decoder core 120 initially only provides MDCT coefficients. Moreover, the filter 130 can optionally derive the spectral values from a time domain representation of the decoded audio representation 120, for example, by a Fourier transform or an MDCT transform or the like (e.g., a short-time Fourier transform STFT).

[0112] The scaling value determination 134 derives scaling values 136 from a plurality of spectral values of the decoded audio representation (e.g., derived from the decoded audio representation). For example, the scaling value determination 134 can comprise a neural network or a machine learning structure that receives the spectral values 132 and derives the scaling values 136. Moreover, the spectral values of the enhanced audio representation 112 can be obtained by scaling spectral values of the decoded audio representation (which can be equal or different from the spectral values used by the scaling value determination 134) according to the scaling values 136. For example, the scaling values 136 can define a scaling of the spectral values in different frequency bins or frequency ranges. Moreover, it should be noted that the scaling 136 can operate on complex-valued spectral values or on real-valued spectral values (e.g., magnitude values or amplitude values or norm values).

[0113] Thus, when an appropriate determination of the scaling value 136 is used based on the spectral values 132 of the decoded audio representation, the scaling 138 can counteract a deterioration of the audio quality caused by the lossy encoding used to provide the encoded audio representation 110.

[0114] For example, the scaling 138 can reduce quantization noise, e.g., by selectively attenuating spectral bins or spectral ranges that comprise high quantization noise. Alternatively or in addition, the scaling 138 can also cause a smoothing of the spectrum over time and / or over frequency, which can also help to reduce quantization noise and / or to improve the perceptual impression.

[0115] However, it should be noted that the audio decoder 100 according to Figure 1 may optionally be supplemented by any of the features, functions, and details disclosed herein, individually and in combination.

[0116] 2) According to Figure 2 device

[0117] Figure 2 A schematic block diagram of an apparatus 200 for determining a set of values defining a filter (e.g., a neural network-based filter, or a filter based on another machine learning structure) is shown.

[0118] The apparatus 200 according to Figure 2 is configured to receive a decoded audio representation 210 and to provide a set of values defining a filter 212 based on the decoded audio representation 210, wherein the set of values defining the filter 212 can for example comprise coefficients of a neural network or coefficients of another machine learning structure. Optionally, the apparatus 200 can receive a target scaling value 214 and / or target spectral information 216. However, the apparatus 200 can optionally generate the target scaling value and / or the target spectral information 216 itself.

[0119] It should be noted that the target scaling value can for example describe a scaling value that brings the decoded audio representation 210 closer to (or closer to) an ideal (undistorted) state. For example, the target scaling value can be determined based on knowledge of a reference audio representation from which the decoded audio representation 210 is derived by encoding and decoding. For example, it can be derived from knowledge of spectral values of the reference audio representation and knowledge of spectral values of the decoded audio representation, wherein the scaling brings an enhanced audio representation (obtained using the scaling based on the spectral values of the decoded audio representation) closer to the reference audio representation.

[0120] Furthermore, the target spectral information 216 can for example be based on knowledge of a reference audio representation from which the decoded audio representation is derived by encoding and decoding. For example, the target spectral information can take the form of spectral values of the reference audio representation.

[0121] AsFigure 2 As can be seen, the apparatus 200 can optionally comprise a spectrum value determination, wherein the spectrum values of the decoded audio representation 210 are derived from the decoded audio representation 210. The spectrum value determination is designated 220, while the spectrum values of the decoded audio representation are designated 222. It should be noted, however, that the spectrum value determination 220 should be considered optional, since the decoded audio representation 210 can be provided directly in the form of spectrum values.

[0122] The apparatus 200 further comprises a determination 230 of a set of values defining a filter. The determination 230 can receive or obtain the spectrum values 222 of the decoded audio representation and provide the set of values 212 defining the filter based on the spectrum values 222. The determination 230 can optionally use the target scaling values 214 and / or the target spectrum information 216.

[0123] With respect to the functionality of the apparatus 200, it should be noted that the apparatus 200 is configured to obtain spectrum values 222 of the decoded audio representation associated with different frequency segments or frequency ranges. Further, the determination 230 can be configured to determine the set of values 212 defining the characteristics of the filter such that the scaling values provided by the filter based on the spectrum values 222 of the decoded audio representation associated with different frequency segments or frequency ranges approximate the target scaling values (e.g., the target scaling values 214). As mentioned above, the target scaling values can be computed based on a comparison of a desired enhanced audio representation and the decoded audio representation, wherein the desired enhanced audio representation can correspond to the aforementioned reference audio representation. In other words, the determination 230 can determine and / or optimize the set of values defining the characteristics of the filter (e.g., a set of coefficients of a neural network, or a set of coefficients of another machine learning structure), e.g., a neural network-based filter, or a filter based on another machine learning structure, such that the filter provides scaling values based on the spectrum values of the decoded audio representation that approximate the target scaling values 214. The determination of the set of values 214 defining the filter can be done using a single pass forward computation, but can typically be performed using an iterative optimization. However, any known training procedure for neural networks or for computer learning structures can be used.

[0124] Alternatively, the determining 230 of the set of values 212 defining the filter can be configured to determine the set of values 212 defining the characteristics of the filter such that a spectrum obtained by the filter based on spectral values of the decoded audio representation associated with different frequency bands or frequency ranges and using the scaling values obtained based on the decoded audio representation approximates a target spectrum (which can be described, for example, by the target spectrum information 216). In other words, the determining 230 can select the set of values 212 defining the filter such that the filtered version of the spectral values of the decoded audio representation 210 approximates the spectral values described by the target spectrum information 216. In summary, the apparatus 200 can determine the set of values 212 defining the filter such that the filter at least partly brings the spectral values of the decoded audio representation to the “ideal” or “reference” or “target” spectral values. For this purpose, the apparatus typically uses the decoded audio representation representing different audio content. By determining the set of values 212 defining the filter based on different audio content (or different types of audio content), the set of values 212 defining the filter can be selected such that the filter performs quite well for audio content different from the reference audio content used to train the set of values 212 defining the filter.

[0125] Thus, it can be achieved that the set of values 212 defining the filter is very well suited for enhancing the decoded audio representation obtained in an audio decoder (e.g., in the audio decoder 100 according to Figure 1 In other words, the set of values 212 defining the filter can be used, for example, in the audio decoder 100 for defining the operation of the scaling value determination 134 (and thus of the filter 130).

[0126] However, it should be noted that the apparatus 200 according to Figure 2 may optionally be supplemented by any of the features, functions, and details described herein, individually and in combination.

[0127] 3) Audio decoder 300 according to Figure 3

[0128] Fig. 3 shows a schematic block diagram of an audio decoder 300 according to a further embodiment of the present application. The audio decoder 300 is configured to receive an encoded audio representation 310 (which can correspond to the encoded audio representation 110) and to provide, based on the encoded audio representation 310, an enhanced audio representation 312 (which can correspond to the enhanced audio representation 112). The audio decoder 300 comprises a decoder core 320 (which can correspond to the decoder core 120). The decoder core 320 provides a decoded audio representation 322 (which can correspond to the decoded audio representation 122) based on the encoded audio representation 310. The decoded audio representation can be in a time domain representation, but also in a spectral domain representation.

[0129] Optionally, the audio decoder 300 can comprise a conversion 324 which can receive the decoded audio representation 322 and provide a spectral domain representation 326 based on the decoded audio representation 322. For example, the conversion 324 can be useful if the decoded audio representation does not take the form of spectral values associated with different frequency bins or frequency ranges. For example, if the decoded audio representation 322 is in a time domain representation, the conversion 324 can convert the decoded audio representation 322 into a plurality of spectral values. However, in case the decoder core 320 does not provide spectral values which can be used by subsequent processing stages, the conversion 324 can also perform a conversion from a first type of spectral domain representation to a second type of spectral domain representation. The spectral domain representation 326 may, for example, comprise spectral values 132 as shown in the audio decoder 100 of Figure 1

[0130] Furthermore, the audio decoder 300 comprises a scaling value determination 334 which, for example, comprises an absolute value determination 360, a logarithm computation 370, and a neural network or machine learning structure 380. The scaling value determination 334 provides a scaling value 336 based on the spectral values 326 which can correspond to the spectral values 132.

[0131] The audio decoder 300 further comprises a scaling 338 which can correspond to the scaling 138. In the scaling, the spectral values of the decoded audio representation or a pre-processed version thereof are scaled according to the scaling value 336 provided by the neural network / machine learning structure 380. Thus, the scaling 338 provides an enhanced audio representation.

[0132] The scaling value determination 334 and the scaling 338 can be considered as filters or “post-filters”.

[0133] In the following, some further details will be described.

[0134] The scaling value determination 334 comprises an absolute value determination 360. The absolute value determination 360 can receive the spectral domain representation 326 of the decoded audio representation, for example The absolute value determination 360 can then provide an absolute value 362 of the spectral domain representation 326 of the decoded audio representation. The absolute value 362 may, for example, be specified as

[0135] The scaling value determination further comprises a logarithm computation 370 which receives the absolute values 362 of the spectral domain representation of the decoded audio representation (e.g. a plurality of absolute values of spectral values) and provides a logarithmized absolute value 372 of the spectral domain representation of the decoded audio representation based on the absolute values 362. For example, the logarithmized absolute value 372 can be specified as

[0136] ​It should be noted that the absolute value determination 360 can for example determine the absolute value or amplitude value or norm value of a plurality of spectral values of the spectral domain representation 326 such that for example the sign or phase of the spectral values is removed. The logarithm calculation is for example a calculation of the common logarithm (base 10) or the natural logarithm or any other possibly suitable logarithm. Further, it should be noted that the logarithm calculation can optionally be replaced by any other calculation reducing the dynamic range of the spectral values 362. Further, it should be known that the logarithm calculation 370 can comprise a limiting of negative values and / or positive values such that the logarithmized absolute values 372 can be limited to a reasonable range of values.

[0137] The scaling value determination 334 further comprises a neural network or machine learning structure 380 receiving the logarithmized absolute values 372 and providing the scaling values 332 based on the logarithmized absolute values 372. The neural network or machine learning structure 380 can for example be parameterized by a set of values 382 defining characteristics of the filter. For example, the set of values can comprise coefficients of the machine learning structure or coefficients of the neural network. For example, the set of values 382 can comprise branch weights of the neural network and optionally also parameters of an activation function. The set of values 382 can for example be determined by the apparatus 200 and the set of values 382 can for example correspond to the set of values 212.

[0138] Further, the neural network or machine learning structure 380 can also optionally comprise logarithmized absolute values of spectral domain representations of decoded audio representations of one or more frames before the current frame and / or one or more frames after the current frame. In other words, the neural network or machine learning structure 380 can not only use the logarithmized absolute values of spectral values associated with the current processed frame (for which the scaling values are applied) but also take into account the logarithmized absolute values of spectral values of one or more previous frames and / or one or more subsequent frames. Thus, the scaling values associated with a given (currently processed) frame can be based on spectral values of the given (currently processed) frame and also on spectral values of one or more previous frames and / or one or more subsequent frames.

[0139] For example, the logarithmized absolute values of spectral domain representations of decoded audio representations (designated as 372) can be applied to the input (e.g. input neurons) of the neural network or machine learning structure 380. The scaling values 336 can be provided by the output (e.g. by output neurons) of the neural network or machine learning structure 380. Further, the neural network or machine learning structure can perform the processing in accordance with the set of values 382 defining characteristics of the filter.

[0140] The scaling 338 can receive the scaling values 336 (which can also be designated as "masking values" and which may, for example, be designated as M(k, n)) and also receive spectral values of the decoded audio representation, or pre-processed spectral values of a spectral domain representation of the decoded audio representation. For example, the spectral values that are input into the scaling 338 and scaled according to the scaling values 336 can be based on the spectral domain representation 326 or can be based on the absolute values 362, wherein a pre-processing can optionally be applied prior to performing the scaling 338. The pre-processing may, for example, comprise a filtering, for example, in the form of a fixed scaling or in the form of a scaling determined by the side information of the encoded audio information. However, the pre-processing can also be fixed and can be independent of the side information of the encoded audio representation. Moreover, it should be noted that the spectral values that are input into the scaling 338 and scaled using the scaling values 336 do not necessarily need to be identical to the spectral values used for deriving the scaling values 336.

[0141] Thus, the scaling 338 may, for example, multiply the spectral values input into the scaling 338 with the scaling values, wherein different scaling values are associated with different frequency bands or frequency ranges. Thus, the enhanced audio representation 312 is obtained, wherein the enhanced audio representation may, for example, comprise a scaled spectral domain representation (e.g., or a scaled absolute value of such a spectral domain representation (e.g. Thus, the scaling 338 may, for example, be performed using a simple multiplication between the spectral values associated with the decoded audio representation 322 and the associated scaling values provided by the neural network or the machine learning structure 380.

[0142] In summary, the apparatus 300 provides an enhanced audio representation 312 based on the encoded audio representation 310, wherein a scaling 338 is applied to spectral values based on the decoded audio representation 322 provided by the decoder core 320. The scaling values 336 used in the scaling 338 are provided by a neural network or by a machine learning structure, wherein the input signal of the neural network or the machine learning structure 380 is preferably obtained by logarithmizing absolute values based on the spectral values of the decoded audio representation 322. However, by a suitable selection of the set of values 382 defining the characteristics of the filter, the neural network or the machine learning structure can provide scaling values in such a way that the scaling 338 improves the auditory impression of the enhanced audio representation when compared to the decoded audio representation.

[0143] Moreover, it should be noted that the audio decoder 300 can optionally be supplemented by any of the features, functions and details described herein.

[0144] 4) Apparatus according to Figure 4

[0145] Fig. 4 shows a schematic block diagram of an apparatus 400 for determining a set of values of a feature defining a filter, e.g. coefficients of a neural network, or coefficients of another machine learning structure. The apparatus 400 is configured to receive a training audio representation 410 and to provide a set of values of a feature defining a filter 412 based on the training audio representation 410. It should be noted that the training audio representation 410 may, for example, comprise different audio content for determining the set of values 412.

[0146] The apparatus 400 comprises an audio encoder 420 configured to encode the training audio representation 410, thereby obtaining an encoded training audio representation 422. The apparatus 400 further comprises a decoder core 430 receiving the encoded training audio representation 422 and providing a decoded audio representation 432 based on the training audio representation 422. It should be noted that the decoder core 420 may, for example, be identical to the decoder core 320 and the decoder core 120. The decoded audio representation 432 may, for example, also correspond to the decoded audio representation 210.

[0147] The apparatus 400 further optionally comprises a conversion 442 converting the decoded audio representation 432 based on the training audio representation 410 to a spectral domain representation 446. The conversion 442 may, for example, correspond to the conversion 324 and the spectral domain representation 446 may, for example, correspond to the spectral domain representation 326. The apparatus 400 further comprises an absolute value determination 460 receiving the spectral domain representation 446 and providing an absolute value of the spectral domain representation 462 based on the spectral domain representation 446. The absolute value determination 460 may, for example, correspond to the absolute value determination 360. The apparatus 400 further comprises a logarithm computation 470 receiving the absolute value of the spectral domain representation 462 and providing a logarithmized absolute value of the spectral domain representation of the decoded audio representation 472 based on the absolute value 462. The logarithm computation 470 may, for example, correspond to the logarithm computation 370.

[0148] Further, the apparatus 400 comprises a neural network or machine learning structure 480 corresponding to the neural network or machine learning structure 380. However, the coefficients of the machine learning structure or neural network 480 designated as 482 are provided by a neural network training / machine learning training 490. It should be noted here that the neural network / machine learning structure 480 provides a scaling value to the neural network training / machine learning training 490, which the neural network / machine learning structure derives based on the logarithmized absolute value 372.

[0149] The apparatus 400 further comprises a target scaling value computation 492, which is also designated as "ratio mask computation". For example, the target scaling value computation 492 receives the absolute values 462 of the spectral domain representation of the training audio representation 410 and the decoded audio representation 432. Thus, the target scaling value computation 492 provides target scaling value information 494, which describes the desired scaling values that should be provided by the neural network / machine learning structure 480. Thus, the neural network training / machine learning training 490 compares the scaling values 484 provided by the neural network / machine learning structure 480 with the target scaling values 494 provided by the target scaling value computation 492 and adjusts the values 482 (i.e. the coefficients of the machine learning structure or the neural network) in order to reduce (or minimize) the deviation between the scaling values 484 and the target scaling values 494.

[0150] In the following, an overview of the functionality of the apparatus 400 will be provided. By encoding and decoding the training audio representation (which may, for example, comprise different audio content) in the audio encoder 420 and the audio decoder 430, a decoded audio representation 432 is obtained, which typically comprises some deterioration (compared to the training audio representation) due to the loss in the lossy encoding. The target scaling value computation 492 determines which scaling (e.g. which scaling values) should be applied to the spectral values of the decoded audio representation 432 such that the scaled spectral values of the decoded audio representation 432 are very close to the spectral values of the training audio representation. It is assumed that by applying a scaling to the spectral values of the decoded audio representation 432, at least partly, the artifacts introduced by the lossy encoding can be compensated. Thus, the neural network or the machine learning structure 480 is trained by the neural network training / machine learning training such that the scaling values 482 provided by the neural network / machine learning structure 480 based on the decoded audio representation 432 are close to the target scaling values 494. The optional conversion 442, the absolute value determination 460 and the logarithm computation 470 only constitute (optional) pre-processing steps to derive the input values 472 for the neural network or the machine learning structure 480 (which are the logarithmized absolute values of the spectral values of the decoded audio representation).

[0151] The neural network training / machine learning training 490 can use an appropriate learning mechanism (e.g. an optimization process) to adjust the coefficients 482 of the machine learning structure or the neural network such that the difference (e.g. a weighted difference) between the scaling values 484 and the target scaling values 494 is minimized or below a threshold or at least reduced.

[0152] Thus, the coefficients 482 of the machine learning structure or the neural network (or, in general, a set of values defining the characteristics of a filter) are provided by the apparatus 400. These values can be used in the filter 130 (to adjust the scaling value determination 134) or in the apparatus 300 (to adjust the neural network / machine learning structure 380).

[0153] It should be noted, however, that the apparatus 400 can optionally be supplemented by any of the features, functions and details described herein.

[0154] 5. Apparatus according to Figure 5

[0155] Fig. 5 shows a schematic block diagram of an apparatus 500 for determining a set 512 of values defining a filter, wherein the values 512 can be, for example, coefficients of a machine learning structure or neural network.

[0156] It should be noted that the apparatus 500 is similar to the apparatus 400, so that the same features, functions and details will not be outlined again. Instead, reference is made to the above description.

[0157] The apparatus 500 receives a training audio representation 510, which can, for example, correspond to the training audio representation 410. The apparatus 500 comprises an audio encoder 520, which corresponds to the audio encoder 420 and provides an encoded training audio representation 522, which corresponds to the encoded training audio representation 422. The apparatus 500 further comprises a decoder core 530, which corresponds to the decoder core 430 and provides a decoded audio representation 532.

[0158] The apparatus 500 optionally comprises a transformation 542, which corresponds to the transformation 442 and provides a spectral domain representation of the decoded audio representation 552, for example, in the form of spectral values. The spectral domain representation is designated as 546 and corresponds to the spectral domain representation 446. Further, the apparatus 500 comprises an absolute value determination 560, which corresponds to the absolute value determination 460. The apparatus 500 further comprises a logarithm computation 570, which corresponds to the logarithm computation 470. Further, the apparatus 500 comprises a neural network or machine learning structure 580, which corresponds to the machine learning structure 480. However, the apparatus 500 further comprises a scaling 590, which is configured to receive the spectral values 546 of the decoded audio representation or the absolute values 562 of the spectral values of the decoded audio representation. The scaling further receives a scaling value 584 provided by the neural network 580. Thus, the scaling 590 scales the spectral values of the decoded audio representation or the absolute values of the spectral values of the audio representation, thereby obtaining an enhanced audio representation 592. The enhanced audio representation 592 can, for example, comprise scaled spectral values (e.g. ) or scaled absolute values of spectral values (e.g. ). In principle, the enhanced audio representation 592 can correspond to the enhanced audio representation 112 provided by the apparatus 100 and the enhanced audio representation 312 provided by the apparatus 300. In this regard, the functionality of the apparatus 500 can correspond to the functionality of the apparatus 100 and / or the functionality of the apparatus 300, except that the coefficients of the neural network or machine learning structure 580, designated as 594, are adjusted by the neural network training / machine learning training 596. For example, the neural network training / machine learning training 596 can receive the training audio representation 510 and can further receive the enhanced audio representation 592, and can adjust the coefficients 594 such that the enhanced audio representation 592 approximates the training audio representation.

[0159] It is noted here that, if the enhanced audio representation 592 approximates the training audio representation 510 with good accuracy, then the signal deterioration caused by the lossy coding is at least partially compensated by the scaling 590. In other words, the neural network training 596 can, for example, determine a (weighted) difference between the training audio representation 510 and the enhanced audio representation 592, and adjust the coefficients 594 of the machine learning structure or neural network 580 to reduce or minimize this difference. The adjustment of the coefficients 594 can be performed, for example, in an iterative process.

[0160] It can thus be achieved that the coefficients 594 of the neural network or machine learning structure 580 are adapted such that, in normal operation, the machine learning structure or neural network 380 using the determined coefficients 594 can provide the scaling values 336 which result in a good quality enhanced audio representation 312.

[0161] In other words, the coefficients 482, 594 of the neural network or machine learning structure 480 or 580 can be used in the neural network 380 of the apparatus 300, and in this case, it can be expected that the apparatus 300 provides a high quality enhanced audio representation 312. Of course, this functionality is based on the assumption that the neural network / machine learning structure 380 is similar or identical to the neural network / machine learning structure 480 or 580.

[0162] Further, it is noted that the coefficients 482, 412 or the coefficients 594, 512 can also be used in the scaling value determination 134 of the audio decoder 100.

[0163] Further, it is noted that the apparatus 500 can optionally be supplemented by any of the features, functionalities and details described herein, individually and in combination.

[0164] 6). Details and embodiments

[0165] In the following, some considerations underlying the present application will be discussed and several solutions will be described. In particular, a number of details will be disclosed, which can optionally be introduced into any of the embodiments disclosed herein.

[0166] 6.1 Problem formulation

[0167] 6.1.1 Ideal Ratio Mask (IRM)

[0168] From a very simple mathematical point of view, the encoded speech The decoded speech (e.g., provided by a decoder core (e.g., decoder core 120 or decoder core 320 or decoder core 430 or decoder core 530)) can be described as:

[0169]

[0170] where x(n) is the input of the encoder (e.g., input of the audio encoder 410, 510) and δ(n) is the quantization noise. Since ACELP uses a perceptual model during the quantization process, the quantization noise δ(n) is correlated with the input speech. This correlation property of the quantization noise makes the post-filtering problem unique for the speech enhancement problem, which assumes that the noise is uncorrelated. To reduce the quantization noise, a real-valued mask is estimated for each time-frequency bin and this mask is multiplied with the amplitude of the encoded speech for this time-frequency bin.

[0171]

[0172] where M(k, n) is the real-valued mask, is the amplitude of the encoded speech, is the amplitude of the enhanced speech, k is the frequency index, and n is the time index. If the mask is ideal (e.g., if the scaling values M(k, n) are ideal), the clean speech can be reconstructed from the encoded speech.

[0173]

[0174] where |x(k, n)| is the amplitude of the clean speech.

[0175] Comparing equation 2 and equation 3, the ideal ratio mask (IRM) (e.g., ideal values of the scaling values M(k, n)) is obtained and given by

[0176]

[0177] where γ is a very small constant factor to prevent division by zero. Since the amplitude values lie in the range [0, ∞], the values of the IRM also lie in the range [0, ∞].

[0178] In other words, for example, the enhanced audio representation may be based on the decoded audio is derived using scaling, wherein the scaling factor can be described by M(k, n). Further, for example, the scaling factor M(k, n) can be derived from the decoded audio representation, since there is usually a correlation between the noise (at least partially compensated by using the scaling factor M(k, n)) and the decoded audio representation . For example, the scaling given in equation (2) can be performed by the scaling 138, wherein the scaling value determination 134 can, for example, provide the scaling value M(k, n) which will be close to the ideal scaling vector IRM(k, n) as described in equation (4).

[0179] Thus, it is desired that the scaling value determination 134 determines a scaling value which is close to IRM(k, n).

[0180] This can be achieved, for example, by a proper design of the scaling value determination 134 or the scaling value determination 334, wherein, for example, the coefficients of a machine learning structure or neural network used to implement the block 380 can be determined as outlined in the following.

[0181] 6.1.2 MMSE optimization

[0182] For example, two different types of minimum mean square error (MMSE) optimization can be used to train a neural network (e.g., the neural network 380): mask approximation (MA) (e.g., as shown in Fig. 4) and signal approximation (SA)

[10] (e.g., as shown in Fig. 5). The MA optimization method attempts to minimize the mean square error (MSE) between the target mask (e.g., the target scaling value) and the estimated mask (e.g., the scaling value 484 provided by the neural network).

[0183]

[0184] wherein IRM(k, n) is the target mask and M(k, n) is the estimated mask.

[0185] The SA optimization method attempts to minimize the mean square error (MSE) between the target amplitude spectrum |X(k, n) (e.g., the amplitude spectrum of the training audio representation 510) and the enhanced amplitude spectrum (e.g., the amplitude spectrum of the enhanced audio representation 592).

[0186]

[0187] wherein the enhanced amplitude spectrum is given by equation 2.

[0188] In other words, the neural network used in the scaling value determination 134 or the scaling value determination 334 can be trained, for example as shown in Figures 4 and 5. As can be seen from Figure 4, the neural network training / machine learning training 490 optimizes the neural network coefficients or machine learning structure coefficients 482 according to the criteria defined in equation (5).

[0189] As shown in Figure 5, the neural network training / machine learning training 596 optimizes the neural network coefficients / machine learning structure coefficients 594 according to the criteria shown in equation (6).

[0190] 6.1.3 Analysis of the mask values

[0191] In most of the proposed mask-based speech enhancement and dereverberation methods, the mask values are bound to 1 [9]

[10] . This is because, generally, if the mask values are not bound to 1, estimation errors can cause amplification of noise or tonality

[15] . Therefore, these methods use sigmoid as the output activation in order to bind the mask values to 1.

[0192] Table 1 shows the percentage of mask values lying in the interval (0, 1) for different signal-to-noise ratios (SNRs). These mask values are computed by adding white noise of different SNRs to clean speech. We can infer from Table 1 that most of the mask values lie in the interval [0, 1], so binding the mask values to 1 does not have a detrimental effect on the neural network-based speech enhancement system.

[0193] We then compute the distribution of the mask values at the lower three bitrates (6.65 kbps, 8.85 kbps, and 12.65 kbps) of AMR-WB. Table 2 shows the computed distribution. One major difference from Table 1 is the percentage of mask values lying in the range [0, 1]. While 39% of the values lie in this range at 6.65 kbps, this value increases to 44% at 12.65 kbps. Almost 30%-36% of the mask values lie in the range [1, 2]. Almost 95% of the mask values lie in the range [0, 5]. Therefore, for the post-filtering problem, we cannot simply bind the mask values to 1. This prevents us from using sigmoid activation (or a simple non-scaled sigmoid activation) at the output layer.

[0194] In other words, it has been found that it is advantageous to use a mask value (also designated as scaling value) larger than 1 in embodiments according to the present application. Furthermore, it has been found that it is advantageous to limit the mask value or scaling value to a predetermined value which should be larger than 1 and which may, for example, be in the region between 1 and 10 or in the region between 1.5 and 10. By limiting the mask value or scaling value, an over-scaling which can lead to artifacts can be avoided. A scaling value in the appropriate range can be achieved, for example, by using a scaled sigmoid activation in the output layer of the neural network or by using a (e.g. rectified) linear activation function as output layer of the neural network.

[0195] 6.2 Experimental setup

[0196] In the following, some details regarding the experimental setup will be described. However, it should be noted that the feature functions and details described herein can optionally be employed into any of the embodiments disclosed herein.

[0197] Our proposed post-filter computes (e.g. in block 324) a short-time Fourier transform (STFT) of frames with a 50% overlap rate (8 ms) at a 16 kHz sampling rate and a length of 16 ms. Before computing a fast Fourier transform (FFT) of length 256 resulting in 129 frequency bins (e.g. spatial domain representation 326), the time frames are windowed with a Hanning window. From the FFT, log-magnitude values are computed to compress the very large dynamic range of the magnitude values (e.g. log-arithmized absolute values 372). Since speech is time-dependent, we use context frames around the time frame under processing (e.g. designated as 373). We tested our proposed model in two cases: a) only past context frames were used, and b) both past and future context frames were used. This was done because future context frames would increase the delay of the proposed post-filter, and we wanted to test the benefit of using future context frames. When only past context frames were considered, we chose a context window of 3, resulting in a delay of only one frame (16 ms). When both past and future context frames were considered, the delay of the proposed post-filter was 4 frames (64 ms).

[0198] When testing with only the past 3 context frames and the current processing frame, the input feature dimension of our proposed neural network (e.g. values 373 and 373) was 516 (4*129). When testing with past and future context frames, the input feature dimension was 903 (7*129). The input features (e.g. values 372 and 373) were normalized to zero mean and unit variance. However, neither the target real-valued mask (e.g. value 494) nor the target magnitude spectrum of the un-coded speech (e.g. magnitude values 410) were normalized.

[0199] Figure 6 An FCNN 600 is shown, which is trained to learn a mapping function f between the log-magnitudes and the real-valued masks θ .

[0200]

[0201] An FCNN is a simple neural network with an input layer 610, one or more hidden layers 612a to 612d, and an output layer 614. We implemented the FCNN in python using Keras

[16] with Tensorflow

[17] as the backend. In our experiments, we have used 4 hidden layers with 2048 units each. All 4 hidden layers use a rectified linear unit (ReLU) as the activation function

[18] . The output of the hidden layers is normalized using batch normalization

[19] . To prevent overfitting, we set the dropout

[20] to 0.2. To train our FCNN, we used the Adam optimizer

[21] with a learning rate of 0.01 and a batch size of 32.

[0202] The output layer 614 has a dimension of 129. Since our FCNN estimates relative (or real-valued) masks, and these masks can be any value between [0, ∞], we have tested both with bound and unbound mask values. When the mask values are unbound, we use a ReLU activation in the output layer. When the mask values are bound, we use either a bound ReLU activation or use a sigmoid function, and scale the output of the sigmoid activation function by a certain scaling factor N.

[0203] To train our FCNN, we used two loss functions (MSE MA and MSE SA ) as defined in Section 6.1.2. When either the bound ReLU or unbound ReLU is used as the output layer activation, the clip norm is used to ensure the convergence of the model.

[0204] When using either the bound or unbound ReLU, the gradient of the output layer is:

[0205]

[0206] where tar is the magnitude spectrum (e.g., of the audio representation 510) or IRM (e.g., values 494), out is the enhanced magnitude (e.g., values 542) or estimated mask (e.g., values 484) (taking any value between 0 and the threshold), and h is the output of the hidden unit (as input to the output unit). When using a bound ReLU, equation 8 is zero outside the bound values.

[0207] When using a scaled sigmoid, the gradient of the output layer is:

[0208]

[0209] where tar is the magnitude spectrum or IRM (e.g., values 494), out is the enhanced magnitude or estimated mask Mest (taking any value between 0 and 1), and h is the output of the hidden unit (as input to the output unit).

[0210] For our training, validation, and testing, we used the NTT database

[22] . We also performed a cross-database test on the TIMIT database

[23] to confirm the independence of the model from the training database. Both the NTT and TIMIT databases are clean speech databases. The TIMIT database consists of monophonic speech files sampled at 16 kHz. The NTT database consists of stereo speech files sampled at 48 kHz. To obtain monophonic speech files at 16 kHz, we performed passive downmixing and resampling on the NTT database. The NTT database consists of 3960 files, of which 3612 files are used for training, 198 files are used for validation, and 150 files are used for testing. The NT database consists of both male and female speakers and also includes languages such as American and British English, German, Chinese, French, and Japanese.

[0211] Time-domain enhanced speech is obtained using an inverse short-time Fourier transform (iSTFT). The iSTFT uses the phase of the encoded speech without any processing.

[0212] In summary, the scaled value determination 134 or the neural network 380 is implemented in embodiments according to the present application using a fully connected neural network 600 as shown in Figure 6 Fig. 6. Furthermore, the neural network 600 can be trained by the apparatus 200 or the apparatus 400 or the apparatus 500.

[0213] It can be seen that the neural network 600 receives in its input layer 610 log- scaled magnitude values (e.g., log-scaled absolute values 132, 372, 472, 572 of spectral values). For example, log-scaled absolute values of spectral values of a current processing frame and one or more previous frames and one or more subsequent frames can be received at the input layer 610. The input layer can, for example, receive log-scaled absolute values of spectral values. The values received by the input layer can then be forwarded in a scaled manner to the artificial neurons of the first hidden layer 612a. The scaling of the input values of the input layer 612 can, for example, be defined by a set of values defining a characteristic of a filter. Subsequently, the artificial neurons of the first hidden layer 612 (which can be implemented using a non-linear function) provide output values of the first hidden layer 612a. The output values of the first hidden layer 612a can then be provided in a scaled manner to the input of the artificial neurons of the subsequent (second) hidden layer 612b. Again, the scaling is defined by a set of values defining a characteristic of a filter. Additional hidden layers (which comprise similar functionality) can be included. Finally, the output signal of the last hidden layer (e.g., fourth hidden layer 612d) is provided in a scaled manner to the input of the artificial neurons of the output layer 614. The functionality of the artificial neurons of the output layer 614 can, for example, be defined by an output layer activation function. Thus, the output value of the neural network can be determined using an evaluation of the output layer activation function.

[0214] Furthermore, it should be noted that the neural network can be “fully connected”, which means that, for example, all input signals of the neural network can contribute to the input signals of all artificial neurons of the first hidden layer and the output signals of all artificial neurons of a given hidden layer can contribute to the input signals of all artificial neurons of a subsequent hidden layer. However, the actual contribution can be determined by a set of values defining a characteristic of a filter, which is typically determined by the neural network training 490, 596.

[0215] Furthermore, it should be noted that the neural network training 490, 596 can, for example, use gradients as provided in equations (8) and (9) when determining the coefficients of the neural network.

[0216] It should be noted that any of the features, functionalities, and details described in this section can optionally be introduced individually and in combination into any of the embodiments disclosed herein.

[0217] 6.3 Experiments and results

[0218] To estimate the bound values of the mask values, we conducted an oracle experiment. In this experiment, as Figure 7We show that we estimate the IRM and bind the IRM to different thresholds. We use objective measures such as Perceptual Evaluation of Speech Quality (PESQ)

[24]

[25]

[26] and Perceptual Objective Listening Quality Assessment (POLQA)

[27] for our evaluation. From Figure 7 It can be concluded that setting the threshold to 1 is not as good as setting the threshold to 2, 4 or 10. There is a very subtle difference between thresholds 2, 4 and 10. Therefore, we choose to bind our mask values to 2 in additional experiments.

[0219] Furthermore, Figure 8 Average PESQ and POLQA scores are shown that evaluate the performance of the proposed method and the EVS post-processor. It can be seen that the application of the concepts described herein leads to an improved improvement of speech quality for the case that the artificial neural network is trained with signal proximity (e.g. as shown in Fig. 5) and mask proximity (e.g. as shown in Fig. 4).

[0220] 7. Conclusion

[0221] It has been found that the quality of speech encoded at lower bit rates is greatly affected due to high quantization noise. Post-filtering is usually employed at low bit rates to mitigate the effects of quantization noise. In the present disclosure, we propose a real-valued mask based post-filter to improve the quality of decoded speech at lower bit rates. To estimate this real-valued mask, we employ for example a fully connected neural network that operates on normalized log-magnitudes. We test our proposal on Adaptive Multi-Rate Wideband (AMR-WB) codec at 3 lower modes (6.65 kbps, 8.85 kbps and 12.65 kbps). Our experiments show improvements in PESQ, POLQA and subjective listening tests.

[0222] In other words, embodiments according to the present invention relate to the concept of using a fully connected network in the context of speech encoding and / or speech decoding. Embodiments according to the present invention relate to speech enhancement. Embodiments according to the present invention relate to post-filtering. Embodiments according to the present invention relate to the concept of handling quantization noise (or more precisely, reducing quantization noise).

[0223] In embodiments according to the present invention, a CNN (Convolutional Neural Network) is used as a mapping function in the cepstral domain.

[14] proposes a statistics- context-based post-filter in the log-magnitude domain.

[0224] In this contribution, we formulate the problem of enhancing encoded speech as a regression problem. A fully connected neural network (FCNN) is trained to learn a mapping function f θThe estimated real-valued mask is then multiplied with the input amplitudes to enhance the coded speech. We evaluate our contribution to the AMR-WB codec at bitrates of 6.65 kbps, 8.85 kbps and 12.65 kbps. In an embodiment, the postfilter can be used in EVS [4] [3] as our reference postfilter. More details are referred to sections 6.1 and 6.2. It can be seen that favorable PESQ and POLQA scores are obtained using embodiments according to the present application.

[0225] In the following, some additional points will be described.

[0226] According to a first aspect, a mask-based postfilter is used in embodiments according to the present application to improve the quality of coded speech.

[0227] a. The mask is real-valued (or the scaling value is real-valued). It is estimated from input features for each frequency bin by a machine learning algorithm (or neural network)

[0228] b.

[0229] c. Where M est (k, n) is the estimated mask, is an amplitude value of the coded speech, and is the post-processed speech at frequency bin k and time index n

[0230] d. The currently used input features are log-amplitude spectra, but any derivative of the amplitude spectrum can also be used.

[0231] According to a second aspect, a limitation on the mask values or scaling values can optionally exist.

[0232] The estimated mask values lie in the range [0, ∞], for example. To prevent such a large range, a threshold can optionally be set. In traditional speech enhancement algorithms, the mask is bound to 1. In contrast, we bound it to a threshold larger than 1. The threshold is determined by analyzing the mask distribution. A useful threshold can lie between 2 and 10, for example.

[0233] a. Since the estimated mask values are bound to a threshold, for example, and since the threshold is larger than 1, the output layer can be a bound rectified linear unit ReLU or a scaled sigmoid.

[0234] b. When using a mask close to the MMSE (minimum mean square error optimization) approach to optimize the machine learning algorithm, the target mask (e.g. target scaling value) can optionally be modified by setting mask values (e.g. target scaling values) above the threshold in the target mask to 1 or can be set to the threshold.

[0235] According to a third aspect, the machine learning algorithm can be used as a fully connected neural network. Long Short Term Memory (LSTM) can also be used as an alternative.

[0236] a. The fully connected neural network consists for example of 4 hidden layers. Each hidden layer consists for example of 2048 or 2500 rectified linear unit (ReLU) activations.

[0237] b. The input dimension of the fully connected neural network depends on the context frame and the FFT size. The latency of the system also depends on the context frame and the frame size.

[0238] c. The size of the context frame can for example be any value between 3 and 5. For our experiments we used for example 256 (16ms @ 16kHz) as frame size and FFT size. The size of the context frame was set to 3, as the gain was very small beyond 3. We also tested future + past context frames and only past context frames.

[0239] According to a fourth aspect, the fully connected network is trained using the following MMSE (Minimum Mean Square Error optimization): mask proximity and signal proximity.

[0240] a. In mask proximity, the mean square error between the target mask (e.g. target scaling values) and the estimated mask (e.g. scaling values determined using the neural network) is minimized. The target mask is for example modified as shown in (2.b) (e.g. in subsection b of the second aspect).

[0241] b. In signal proximity, the mean square error between the enhanced amplitudes (e.g. enhanced amplitude spectrum 592) and the target amplitudes (e.g. amplitude spectrum of the audio representation 510) is minimized. The enhanced amplitudes are obtained by multiplying the mask estimated from the DNN (e.g. from the neural network) with the encoded amplitudes. The target amplitudes are the unencoded speech amplitudes.

[0242] In summary, the embodiments described herein can optionally be supplemented by any of the points or aspects described here. However, it should be noted that the points and aspects described here can be used individually or in combination and can be introduced individually and in combination into any of the embodiments described herein.

[0243] 8. The method according to Figure 9 of claim 7.

[0244] Figure 9 A schematic block diagram of a method 900 for providing an enhanced audio representation based on an encoded audio representation according to an embodiment of the application is shown.

[0245] The method comprises providing 910 an encoded audio representation

[0246] Further, the method comprises obtaining 920 a plurality of scaling values (M(k, n)) associated with different frequency bands or frequency ranges based on the spectral values of the decoded audio representation associated with the different frequency bands or frequency ranges, and the method comprises scaling 930 the spectral values of the decoded audio signal representation or a pre-processed version thereof using the scaling values (M(k, n)) to obtain an enhanced audio representation

[0247] The method 900 can optionally be supplemented, individually and in combination, by any of the features, functions and details described herein.

[0248] 9. The method according to Figure 10 9. The method according to Figure 10

[0249] Figure 10 A schematic block diagram illustrating a method 1000 for determining a set of values of a feature defining a filter for providing an enhanced audio representation based on a decoded audio representation is shown, according to embodiments of the application

[0250] The method comprises obtaining 1010 spectral values of the decoded audio representation associated with different frequency bands or frequency ranges

[0251] The method further comprises determining 1020 a set of values of a feature defining a filter such that scaling values provided by the filter based on the spectral values of the decoded audio representation associated with different frequency bands or frequency ranges approximate target scaling values.

[0252] Alternatively, the method comprises determining 1030 a set of values of a feature defining a filter such that a spectrum obtained by the filter based on the spectral values of the decoded audio representation associated with different frequency bands or frequency ranges and using scaling values obtained based on the decoded audio representation approximates a target spectrum.

[0253] 10. Alternative implementations

[0254] While some aspects have been described in the context of an apparatus, it is clear that other aspects of corresponding methods are described by respective blocks or features of the corresponding method steps. Similarly, aspects described in the context of method steps also represent descriptions of the corresponding block or item or features of a corresponding apparatus. Some or all of the method steps can be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or electronic circuit. In some embodiments, one or more of the most important method steps can be executed by such an apparatus.

[0255] ​​The inventive encoded audio signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (e.g. the Internet).

[0256] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.

[0257] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0258] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can for example be stored on a machine readable carrier.

[0259] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0260] In other words, an embodiment of the inventive method is, therefore, a computer program for performing one of the methods described herein, when the computer program runs on a computer.

[0261] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.

[0262] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0263] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein.

[0264] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0265] A further embodiment according to the application comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0266] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0267] The apparatus described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0268] The apparatus described herein, or any components of the apparatus described herein, can be implemented at least partially in hardware and / or in software.

[0269] The methods described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0270] The methods described herein, or any components of the apparatus described herein, can be performed by hardware means and / or by software means.

[0271] The above-described embodiments are merely illustrative for the principles of the present application. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the appended patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0272] 11. References

[0273] [1] 3GPP, "Speech codec speech processing functions; Adaptive Multi-Rate- Wideband (AMR-WB) speech codec; Transcoding functions," 3rd Generation Partnership Project (3GPP), TS 26.190, V12 2009. [Online] Available: http: / / www.3gpp.org / ftp / Specs / html-info / 26190.htm

[0274] [2] M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache, Y. Kamamoto, K. Kikuiri, S. Ragot, J. Faure, H. Ehara, V. Rajendran, V. Atti, H. Sung, E. Oh, H. Yuan, and C. Zhu, “Overview of the EVS codec architecture.” IEEE, 2015, pp. 5698-5702.

[0275] [3] 3GPP, “TS 26.445, EVS Codec Detailed Algorithmic Description; 3GPP Technical Specification (Release 12),” 3rd Generation Partnership Project (3GPP), TS 26.445, 1 2014. [Online] Available: http: / / www.3gpp.org / ftp / Specs / html-info / 26445.htm

[0276] [4] T. Vaillancourt, R. Salami, and M. Jelnek, “New post-processing techniques for low bit rate celp codecs,” in ICASSP, 2015.

[0277] [5] J.-H. Chen and A. Gersho, “Adaptive postfiltering for quality en- hancement of coded speech,” vol. 3, no. 1, pp. 59-71, 1995.

[0278] [6] T. Speech Coding with Code-Excited Linear Prediction. Springer, 2017. [Online] Available: http: / / www.springer.com / gp / book / 9783319502021

[0279] [7] K. Han, Y. Wang, D. Wang, W. S. Woods, I. Merks, and T. Zhang, “Learning spectral mapping for speech dereverberation and de-noising.”

[0280] [8] Y. Zhao, D. Wang, I. Merks, and T. Zhang, “Dnn-based enhancement of noisy and reverberant speech,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016.

[0281] [9] Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 22, pp. 1849-1858, 2014.

[0282]

[10] F. Weninger, J. R. Hershey, J. L. Roux, and B. Schuller, “Discriminatively trained recurrent neural networks for single-channel speech separation,” in IEEE Global Conference on Signal and Information Processing (GlobalSIP), 2014.

[0283]

[11] D. S. Williamson and D. Wang, “Time-frequency masking in the complex domain for speech dereverberation and denoising.”

[0284]

[12] Z. Zhao, S. Elshamy, H. Liu, and T. Fingscheidt, “A cnn postprocessor to enhance coded speech,” in 16th International Workshop on Acoustic Signal Enhancement (IWAENC), 2018.

[0285]

[13] Z. Zhao, H. Liu, and T. Fingscheidt, “Convolutional neural networks to enhance coded speech,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 4, pp. 663-678, April 2019.

[0286]

[14] S. Das and T. “Postfiltering using log-magnitudespectrum for speech and audio coding,” in Proc. Interspeech 2018, 2018, pp. 3543-3547. [Online] Available: http: / / dx.doi.org / 10.21437 / Interspeech.2018-1027

[0287]

[15] W Mack, S. Chakrabarty, F.-R. S. Braun, B. Edler, and E. Habets, “Single-channel dereverberation using direct mmse optimization and bidirectional lstm networks,” in Proc. Interspeech 2018, 2018, pp. 1314-1318. [Online] Available: http: / / dx.doi.org / 10.21437 / Interspeech.2018-1296

[0288]

[16] F. Chollet et al., “Keras,” https: / / keras.io, 2015.

[0289]

[17] M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mane', R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Vie'gas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, "TensorFlow: Large-scale machine learning on heterogeneous systems," 2015, software available from: tensorflow.org. [Online] Available: http: / / tensorflow.org /

[0290]

[18] X. Glorot, A. Bordes, and Y. Bengio, "Deep sparse rectifier neural networks," in International Conference on Artificial Intelligence and Statistics, 2011, p. 315323.

[0291]

[19] S. Ioffe and C. Szegedy, "Batch normalization: Accelerating deep network training by reducing internal covariate shift," in International Conference on Machine Learning, vol. 37, 2015, pp. 448-456.

[0292]

[20] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, "Dropout: A simple way to prevent neural networks from overfitting," J. Mach. Learn. Res., vol. 15, no. 1, pp. 1929-1958, Jan. 2014. [Online] Available: http: / / dl.acm.org / citation.cfm?id=2627435.2670313

[0293]

[21] D. Kingma and J. Ba, "Adam: A method for stochastic optimization," in arXiv preprint arXiv:1412.6980, 2014.

[0294]

[22] NTT-AT, "Super wideband stereo speech database," http: / / www.ntt-at.com / product / wideband speech, accessed: 09.09.2014. [Online] Available: http: / / www.ntt-at.com / product / wideband speech

[0295]

[23] J. S. Garofolo, L. D. Consortium et al., TIMIT: Acoustic-Phonetic Continuous Speech Corpus. Linguistic Data Consortium, 1993.

[0296]

[24] A. Rix, J. Beerends, M. Hollier, and A. Hekstra, "Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs," in 2001 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2001.

[0297]

[25] ITU-T, “P.862.1: Mapping Function for Transforming P.862 Raw Result Scores to MOS-LQO,” (International Telecommunication Union), Tech. Rep. P.862.1, Nov. 2003.

[0298]

[26] ——, “P.862.2: Wideband Extension to Recommendation P.862 for the Assessment of Wideband Telephone Networks and Speech Codecs,” (International Telecommunication Union), Tech. Rep. P.862.2, Nov. 2005.

[0299]

[27] Perceptual objective listening quality assessment (POLQ4), ITU-T Recommendation P.863, 2011. [Online] Available: http: / / www.itu.int / rec / T-REC-P.863 / en

[0300]

[28] Recommendation BS.1534, Method for the subjective assessment of intermediate quality levels of coding systems, ITU-R, 2003.

Claims

1. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420; 520) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), ​​ wherein the filter (130; 360, 370, 380, 338) is configured to obtain the scaling value (136; 336; M(k, n)) such that the scaling value results in a scaling or amplification of one or more spectral values (132; 326) of the decoded audio representation (122; 322; ) or one or more pre-processed spectral values based on the spectral values (132; 326) of the decoded audio representation (122; 322; ).

2. The audio decoder (100; 300) of claim 1, wherein, the filter (130; 360, 370, 380, 338) is adapted to use a configurable processing structure for providing the scaling values (136; 336; M(k,n)), a configuration of the processing structure being based on a machine learning algorithm.

3. The audio decoder (100; 300) of claim 1, wherein, The filter (130; 360, 370, 380, 338) is configured to determine the scaling values (136; 336; M(k, n)) based on the spectral values (132; 326) of the decoded audio representation (122; 322; ) only in a plurality of frequency bands or frequency ranges.

4. The audio decoder (100; 300) of claim 1, wherein, The filter (130; 360, 370, 380, 338) is configured to obtain the amplitude values of the enhanced audio representation according to the following formula : , wherein M(k,n) is a scaling value, wherein k is a frequency index, wherein n is a time index, wherein is an amplitude value of a spectral value of the decoded audio representation; or wherein the filter is configured to obtain a value of the enhanced audio representation according to the following formula : , wherein M(k,n) is a scaling value, wherein k is a frequency index, wherein n is a time index, wherein is a spectral value of the decoded audio representation.

5. The audio decoder (100; 300) of claim 1, wherein The filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges.

6. The audio decoder (100; 300) of claim 5, wherein an input signal (372) of the neural network (380; 600) or of the machine learning structure represents a logarithmic amplitude, an amplitude or a norm of spectral values of the decoded audio representation, the spectral values being associated with different frequency bands or frequency ranges.

7. The audio decoder (100; 300) of claim 5, wherein, an output signal (336) of the neural network (380; 600) or of the machine learning structure represents the scaling values (136; 336; M(k,n)).

8. The audio decoder (100; 300) of claim 5, wherein training the neural network (380; 600) or the machine learning structure to limit, reduce, or minimize a deviation (MSE MA ) between a plurality of target scaling values (494, IRM(k,n)) and a plurality of scaling values (484, M(k,n)) obtained using the neural network (380; 580; 600) or using the machine learning structure.

9. The audio decoder (100; 300) of claim 5, wherein, Train the neural network (380; 600) or the machine learning structure to limit, reduce or minimize the target amplitude spectrum (510), target amplitude spectrum, target absolute spectrum or target norm spectrum ( The deviation between the amplitude spectrum (592), absolute spectrum, or norm spectrum obtained by scaling the processed spectrum and the MSE (mean squared error) is... SA The scaling of the processed spectrum uses a scaling value (584) provided by the neural network (380; 580; 600) or by the machine learning structure.

10. The audio decoder (100; 300) of claim 5, wherein, The neural network (380; 600) or the machine learning structure is trained such that a scaling of one or more spectral values (132; 326) of a spectral decomposition of the decoded audio representation (122; 322; ) or of one or more pre-processed spectral values based on spectral values of a spectral decomposition of the decoded audio representation is located in a range between 0 and a predetermined maximum value.

11. The audio decoder (100; 300) of claim 10, wherein the maximum value being larger than 1.

12. The audio decoder (100; 300) of claim 5, wherein the neural network (380; 600) or the machine learning structure is trained such that the scaling limits to 2, or to 5, or to 10, or to a predetermined value larger than 1 for one or more spectral values of a spectral decomposition of the decoded audio representation, or for one or more pre-processed spectral values based on spectral values of a spectral decomposition of the decoded audio representation.

13. The audio decoder (100; 300) of claim 5, wherein the neural network (380; 600) or the machine learning structure is trained such that the scaling limits to 2, or to 5, or to 10, or to a predetermined value larger than 1.

14. The audio decoder (100; 300) of claim 5, wherein a number of input features of the neural network (380; 600) or of the machine learning structure is at least 2 times larger than a number of output values of the neural network or of the machine learning structure.

15. The audio decoder (100; 300) of claim 5, wherein The filter (130; 360, 370, 380, 338) is configured to normalize input features of the neural network or the machine learning structure to a predetermined mean value and / or a predetermined variance or standard deviation.

16. The audio decoder (100; 300) of claim 1, wherein The neural network (380; 600) comprises an input layer (610), one or more hidden layers (612a-612d), and an output layer (614).

17. The audio decoder (100; 300) of claim 16, wherein, The one or more hidden layers (612a-612d) use a rectified linear unit as an activation function.

18. The audio decoder (100; 300) of claim 16, wherein, The output layer (614) uses a rectified linear unit or a bound rectified linear unit or a sigmoid function as an activation function.

19. The audio decoder (100; 300) of claim 1, wherein, The filter (130; 360, 370, 380, 338) is configured to obtain short-time Fourier transform coefficients representing spectral values of the decoded audio representation, the spectral values being associated with different frequency bins or frequency ranges.

20. The audio decoder (100; 300) of claim 1, wherein The filter (130; 360, 370, 380, 338) is configured to derive a log-magnitude, amplitude, absolute value, or norm value (372) and to determine the scaling values (136; 336; M(k,n)) based on the log-magnitude, amplitude, absolute value, or norm value.

21. The audio decoder (100; 300) of claim 1, wherein, The filters (130; 360, 370, 380, 338) are configured to decode audio representations (122; 322; ) based on the current frame. The spectral values ​​(132; 326) associated with different frequency bands or frequency ranges of the current frame are used to determine a plurality of scaling values ​​(136; 336; M(k,n)) associated with the current frame.

22. The audio decoder (100; 300) of claim 1, wherein, The filter (130; 360, 370, 380, 338) is configured to determine a plurality of scaling values associated with the current frame based on spectral values (132; 326) of a decoded audio representation (122; 322; ) of one or more frames subsequent to the current frame, the spectral values being associated with different frequency bands or frequency ranges.

23. An apparatus (200; 400; 500) for determining a set of values ​​of features of a defining filter (130; 360, 370, 380, 338), the filter being used to provide an enhanced audio representation (112; 312; 322) based on a decoded audio representation (122; 322). ), wherein The apparatus is configured to obtain spectral values (132; 326) of the decoded audio representation (122; 322) associated with different frequency bins or frequency ranges, and wherein the apparatus is configured to determine a set (382; 412; 512) of values defining characteristics of the filter (130; 360, 370, 380, 338) such that scaling values (136; 336; 484; 584) provided by the filter based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges approximate target scaling values (494), or wherein the apparatus is configured to determine a set (382; 412; 512) of values defining characteristics of the filter (130; 360, 370, 380, 338) such that a spectrum obtained by the filter based on spectral values (132; 326) of the decoded audio representation (122; 322) associated with different frequency bins or frequency ranges and using scaling values (136; 336; 484; 584) obtained based on the decoded audio representation (122; 322) approximates a target spectrum (510).

24. The apparatus (200; 400) of claim 23, wherein The apparatus is configured to train a machine learning structure (380; 480; 580) to reduce or minimize a deviation (MSE MA ) between a plurality of target scaling values (494; IRM(k,n)) and a plurality of scaling values (136; 336; 484; 584; M(k,n)) obtained using a neural network based on spectral values (326; 446; 546) associated with different frequency bands or frequency ranges of the decoded audio representation, the machine learning structure (380; 480; 580) being part of the filter (130; 360, 370, 380, 338) and providing scaling values (136; 336; 484; 584; M(k,n) 25. The apparatus (200; 500) of claim 23, wherein, The device is configured to train a machine learning structure (380; 480; 580) to reduce or minimize a target spectrum (510; ) and the spectrum obtained by scaling the processed spectrum (532; 546) (592; The deviation between (MSE) MA The scaling of the processed spectrum (532; 546) uses the scaling value (584) provided by the machine learning structure.

26. The apparatus (200; 400; 500) of claim 23, wherein, The apparatus is configured to train the machine learning structure (380; 480; 580) such that a scaling for spectral values of the decoded audio representation, or for one or more pre-processed spectral values based on spectral values of the decoded audio representation, is located in a range between 0 and 2, or in a range between 0 and 5, or in a range between 0 and 10.

27. The apparatus (200; 400; 500) of claim 23, wherein, The apparatus is configured to train the machine learning structure (380; 480; 580) such that an amplitude scaling for spectral values of the decoded audio representation, or for one or more pre-processed spectral values based on spectral values of the decoded audio representation, is limited to be located in a range between 0 and a predetermined maximum value.

28. The apparatus (200; 400; 500) according to claim 27, wherein The maximum value is greater than 1.

29. A method (900) for providing an enhanced audio representation based on an encoded audio representation, wherein The method comprises providing (910) a decoded audio representation of the encoded audio representation ), wherein the method comprises obtaining (920) a plurality of scaling values (M(k,n)) associated with different frequency bins or frequency ranges based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges, and wherein the method comprises scaling (930) the spectral values of the decoded audio representation (Y(k, n)) or a pre-processed version thereof using the scaling values (M(k, n)) to obtain the enhanced audio representation (Y(k, n)) ​ wherein the scaling value (136; 336; M(k, n)) is obtained such that the scaling value causes a scaling or amplification of one or more spectral values (132; 326) of the decoded audio representation (122; 322; ), or of one or more pre-processed spectral values based on the spectral values (132; 326) of the decoded audio representation (122; 322; ).

30. A method (1000) for determining a set of values of characteristics defining a filter, the filter being used for providing an enhanced audio representation (Y) based on a decoded audio representation (X) ), wherein The method comprises obtaining (1010) spectral values of the decoded audio representation associated with different frequency bins or frequency ranges, and wherein the method comprises determining (1020) a set of values defining characteristics of the filter such that scaling values provided by the filter based on spectral values of the decoded audio representation associated with different frequency bins or frequency ranges approximate target scaling values, or wherein the method comprises determining (1030) a set of values defining characteristics of the filter such that a spectrum obtained by the filter based on spectral values of the decoded audio representation and using scaling values obtained based on the decoded audio representation, the spectral values being associated with different frequency bins or frequency ranges, approximates a target spectrum.

31. A computer program product comprising a computer program for performing the method of claim 29 or 30 when the computer program is run on a computer.

32. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the neural network (380; 600) or the machine learning structure is trained such that a scaling of one or more spectral values (132; 326) of a spectral decomposition of the decoded audio representation (122; 322; ) or one or more pre-processed spectral values based on spectral values of a spectral decomposition of the decoded audio representation is located in a range between 0 and a predetermined maximum value, wherein the maximum value is greater than 1.

33. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the neural network (380; 600) or the machine learning structure is trained such that the scaling restricts the one or more spectral values of the spectral decomposition of the decoded audio representation or the one or more pre-processed spectral values based on the spectral values of the spectral decomposition of the decoded audio representation to a predetermined value greater than 1.

34. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with the different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the neural network (380; 600) or the machine learning structure is trained such that the scaling restricts the one or more spectral values of the spectral decomposition of the decoded audio representation or the one or more pre-processed spectral values based on the spectral values of the spectral decomposition of the decoded audio representation to a predetermined value greater than 1.

35. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420; 520) on the basis of an encoded audio representation (110; 310; 410; 510). ), wherein The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with the different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the filter (130; 360, 370, 380, 338) is configured to normalize input features of the neural network or the machine learning structure to a predetermined mean value and / or a predetermined variance or standard deviation.

36. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with the different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein an input signal (372) of the neural network (380; 600) or the machine learning structure represents a logarithmic amplitude of spectral values of the decoded audio representation, the spectral values being associated with different frequency bins or frequency ranges.

37. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with the different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the neural network (380; 600) comprises an input layer (610), one or more hidden layers (612a-612d), and an output layer (614); wherein the one or more hidden layers (612a-612d) use a rectified linear unit as an activation function.

38. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bins or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with the different frequency bins or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges; wherein the neural network (380; 600) comprises an input layer (610), one or more hidden layers (612a-612d), and an output layer (614); wherein the output layer (614) uses a rectified linear unit or a tied rectified linear unit or a sigmoid function as an activation function.

39. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412; 512) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k, n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) is configured to derive a log amplitude value (372) and determine the scaling values (136; 336; M(k, n)) based on the log amplitude value.

40. An apparatus (200; 400; 500) for determining a set of values defining characteristics of a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; 312) based on a decoded audio representation (122; 322; 322), the filter being defined by a set of values defining characteristics of a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; 312) based on a decoded audio representation (122; 322; 322), ). wherein The filter is configured to use scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ), wherein the apparatus is configured to obtain spectral values (132; 326) of the decoded audio representation (122; 322) associated with different frequency bands or frequency ranges, and wherein the apparatus is configured to determine a set of values (382; 412; 512) defining characteristics of the filter (130; 360, 370, 380, 338) such that scaling values (136; 336; 484; 584) provided by the filter based on spectral values of the decoded audio representation (122; 322) associated with different frequency bands or frequency ranges and associated with different frequency bands or frequency ranges approximate target scaling values (494), or wherein the apparatus is configured to determine a set of values (382; 412; 512) defining characteristics of the filter (130; 360, 370, 380, 338) such that a spectrum obtained by the filter based on spectral values (132; 326) of the decoded audio representation (122; 322) associated with different frequency bands or frequency ranges and using scaling values (136; 336; 484; 584) obtained based on the decoded audio representation (122; 322) approximates a target spectrum (510).

41. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420; 520) on the basis of an encoded audio representation (110; 310; 410; 510). ), wherein The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k, n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) is configured to obtain the scaling value (136; 336; M(k, n)) such that the scaling value results in an amplification of one or more spectral values (132; 326) of the decoded audio representation (122; 322; ) or one or more pre-processed spectral values based on the spectral values (132; 326) of the decoded audio representation (122; 322; ).

42. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412; 512) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k, n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and The filter is configured to use the scaling values ​​(136; 336; M(k,n)) to scale the decoded audio representation ( The spectral values ​​of (122; 312;) or their preprocessed versions are scaled to obtain the enhanced audio representation (122; 312;). ); wherein the filter (130; 360, 370, 380, 338) is configured to obtain the scaling values (136; 336; M(k, n)) such that the scaling values allow both amplification and attenuation by scaling.

43. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412; 512) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k, n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) is configured to obtain the amplitude values of the enhanced audio representation according to the following formula : , wherein M(k, n) is a scaling value, wherein k is a frequency index, wherein n is a time index, wherein is an amplitude value of a spectral value of the decoded audio representation; or wherein the filter is configured to obtain a value of the enhanced audio representation according to the following formula : , wherein M(k, n) is a scaling value, wherein k is a frequency index, wherein n is a time index, wherein M(k, n) is a scaling value, wherein k is a frequency index, wherein n is a time index, wherein is a spectral value of the decoded audio representation.

44. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420; 520) on the basis of an encoded audio representation (110; 310; 410; 510). ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) comprises a neural network (380; 600) or a machine learning structure configured to provide the scaling values (136; 336; M(k,n)) based on a plurality of spectral values (132; 326) describing the decoded audio representation (122; 322; ), the spectral values being associated with different frequency bins or frequency ranges.

45. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420) on the basis of an encoded audio representation (110; 310; 410), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) is configured to obtain short-time Fourier transform coefficients representing spectral values of the decoded audio representation, the spectral values being associated with different frequency bands or frequency ranges.

46. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 420; 520) on the basis of an encoded audio representation (110; 310; 410; 510). ), wherein The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) is configured to derive a logarithmic amplitude, an amplitude, an absolute value or a norm value (372) and determine the scaling values (136; 336; M(k,n)) based on the logarithmic amplitude, the amplitude, the absolute value or the norm value.

47. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412; 512) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) is configured to determine a plurality of scaling values (136; 336; M(k,n)) associated with the current frame based on spectral values (132; 326) of a decoded audio representation (122; 322) of the current frame associated with different frequency bands or frequency ranges and based on spectral values (132; 326) of a decoded audio representation (122; 322) of one or more frames preceding the current frame associated with different frequency bands or frequency ranges. ​ 48. An audio decoder (100; 300) for providing a decoded audio representation (122; 322; 412; 512) on the basis of an encoded audio representation (110; 310; 410; 510), ), wherein, The audio decoder comprises a filter (130; 360, 370, 380, 338) for providing an enhanced audio representation (112; 312; ) of the decoded audio representation (122; 322; ), wherein the filter is configured to obtain a plurality of scaling values (136; 336; M(k,n)) associated with different frequency bands or frequency ranges based on spectral values (132; 326) of the decoded audio representation associated with different frequency bands or frequency ranges, and wherein the filter is configured to scale spectral values of the decoded audio representation (122; 312; 322) or a pre-processed version thereof using the scaling values (136; 336; M(k, n)) to obtain the enhanced audio representation (122; 312; 322), , , wherein the filter (130; 360, 370, 380, 338) is configured to determine a plurality of scaling values associated with a current frame based on spectral values (132; 326) of a decoded audio representation (122; 322) of one or more frames subsequent to the current frame associated with different frequency bands or frequency ranges.