Masking threshhold determinator, quantization step size determinator, audio encoder, methods and computer program applying a comodulation dependant post-masking modeling

WO2026114512A1PCT designated stage Publication Date: 2026-06-04FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV +1

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2024-11-29
Publication Date
2026-06-04

Smart Images

  • Figure EP2024084194_04062026_PF_FP_ABST
    Figure EP2024084194_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A masking threshold determinator for providing a masking threshold information on the basis of an input audio signal is configured to obtain envelope information describing envelopes of different frequency ranges of the input audio signal. The masking threshold determinator is configured to determine a comodulation strength information describing a comodulation between different frequency ranges of the input audio signal. The masking threshold determinator is configured to apply a post-masking modeling to the envelope information, or to a pre-processed version thereof, to obtain a post-masking-processed envelope information, which describes a masking threshold. The masking threshold determinator is configured to adapt a post-masking decay time of the post-masking modeling in dependence on the comodulation strength information. A quantization step size determinator and an encoder, corresponding methods and a computer program are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] a E i

[0002] MASKING THRESHHOLD DETERMINATOR, QUANTIZATION STEP SIZE DETERMINATOR, AUDIO ENCODER, METHODS AND COMPUTER PROGRAM APPLYING A COMODULATION DEPENDANT POST-MASKING MODELING

[0003] TECHNICAL FIELD

[0004] Embodiments of the present inventive concept relate to a masking threshold determinator, a quantization step size determinator, an audio encoder, a method for providing a masking threshold information, a method for providing a quantization step size information, a method for providing an encoded audio representation and a computer program.

[0005] Embodiments according to the invention are related to a perceptual model considering comodulation masking release by post-masking adaptation.

[0006] BACKGROUND OF THE INVENTION

[0007] It has been recognized that in perceptual audio coding, very compact representations maintaining high audio

[0008] quality can be achieved by exploiting masking effects in the human auditory system. Specifically, it has been recognized that the masking of quantization errors by the audio signal itself allows relatively coarse quantization which leads to a reduction of the bit rate for the storage or transmission of the quantized data. As the masking effect varies overtime and frequency depending on the corresponding characteristics of the audio signal, time and frequency dependent quantizer control is advantageous in many cases (and may be required in some cases). It has been found that, for that, it may be advantageous to convert the audio signal to a suitable time-frequency representation, which is often obtained by applying a Modified Discrete Cosine Transform (MDCT). In parallel, the audio signal may, for example, be analyzed by a perceptual model, which estimates the strength of the masking effects, so that quantization can be applied appropriately.

[0009] It has been found that under some circumstances, maskers with equal low frequency modulation even can mutually reduce their masking effects (1, 2). This phenomenon is

[0010] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX jailBBBliiB

[0011] 2

[0012] called comodulation masking release (CMR). It was shown in an experiment with the setup illustrated in Fig. 6.

[0013] Two types of maskers were used in this experiment. The first was stationary band limited noise 650 centered around 1 kHz with the bandwidth as a control parameter. And the second was noise 620 with the same bandwidth, but multiplied by a randomly, relatively slowly changing amplitude modulation function 630. The test signal 640 was a 1 kHz tone adjusted in its level, so that it just started becoming audible, as in usual masking experiments. The result shown in Fig. 7 for the non-modulated masker 720 looks as expected: the masking threshold increases with bandwidth until the critical bandwidth is reached and then remains constant. The result for the modulated masker 730 however looks surprising. At first, the masking threshold increases with bandwidth, but then decreases again. This means, that although the total masker level increases, the masking threshold decreases. Or, in other words, the noise components added by the bandwidth increase rather help for the audibility of the test tone than prevent it.

[0014] For the outcomes of this experiment and other studies some possible explanations are given.

[0015] auditory system analyses coherence of envelope fluctuations detection of disparity in modulation

[0016] suppression of envelope locking

[0017] flanking components guide to the optimum time for detection

[0018] - dip listening

[0019] distribution of neural activities for comodulated components across (frequency) channels

[0020] An approach for considering CMR in a psychoacoustic model for the direct control of an audio encoder was presented in (3). In (4), a model for the perceptual comparison of differently encoded audio data increases the sensitivity in time-frequency regions with high comodulation.

[0021] In view of the above, there is a desire to create a concept which provides an improved compromise between a quality of a determination of a masking threshold, an implementation complexity and possibly an achievable perceptual audio quality.

[0022] FHHS24EM24 FHIIS24EM24-2024346428. DOCX

[0023] C SUMMARY OF THE INVENTION

[0024] Embodiments according to the invention are defined by the independent claims. Optional improvements of these embodiments are defined by the dependent claims.

[0025] In accordance with a first aspect of the present inventive concept, an embodiment creates a masking threshold determinator for providing a masking threshold information (e.g., pfc(m)) on the basis of an input audio signal. The masking threshold determinator is configured to obtain envelope information (e.g., efc(n)) describing envelopes of different, e.g., overlapping or non-overlapping, frequency ranges of the input audio signal. The envelope information can be represented, for example, by a plurality of envelope signals that can be associated with different frequency ranges which can be different frequency bands. In this embodiment, the masking threshold determinator is configured to further determine a comodulation strength information describing a comodulation between different frequency ranges of the input audio signal. The comodulation strength information can, for example, be (or comprise) a plurality of comodulation strength values (e.g., mfc), or can, for example, be (or comprise) a plurality of comodulation coefficients (e.g., ck'), or can, for example, be (or comprise) respective comodulation strength values (e.g., m associated with different respective frequency ranges, or can, for example, be (or comprise) respective comodulation coefficients (e.g., cj associated with different respective frequency ranges. The described comodulation between different frequency ranges of the input audio signal can, for example, be a description of a comodulation of temporal envelopes in neighbouring or distant frequency bands. Also, a description of a comodulation of an envelope signal associated with the currently considered frequency range or having a frequency of range index fc, with one or more other envelope signals associated with one or more different frequency ranges is conceivable. In the current embodiment the masking threshold determinator is further configured to apply a post-masking modelling (e.g., pm) to the envelope information (e.g., efe(n)), or to a pre-processed version thereof, to obtain a post-masking-processed envelope information (e.g., pfe(m)), which describes a masking threshold. This step can, for example, be done in a per frequency-range manner and can, for example, relate to an averaged and / or down-sampled version of the envelope information. The masking threshold here may, for example, serve as a masking threshold information. The masking threshold determinator is configured to adapt a post-masking decay time of the post-masking modelling in dependence on the comodulation strength information. Here, for example, the post masking decay time can be controlled by a decay

[0026] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX SSMIliE

[0027] factor (e.g., dk) (which may, for example, also be considered as an adjustable decay constant) that may be applied in a per-frequency manner. The comodulation strength information can, for example, be represented by a comodulation coefficient (e.g., ck) or comodulation strength value (e.g., mfe( )) that can additionally be associated with a currently considered frequency range, or in dependence on a comodulation coefficient (e.g., ck) associated with a currently considered frequency range. This is done in order to obtain a faster reduction of a post-masking effect in the presence of a comparatively stronger comodulation when compared to the presence of a comparatively weaker comodulation. This embodiment is based on the finding that an adaptation of the post masking decay time of the post masking modeling in dependence on the comodulation strength information is well adapted to characteristics of the human hearing and therefore results in masking threshold values of good quality without requiring excessive computational complexity. For example, the concept allows to have a relatively short decay time of the post masking modeling in the case of a comparatively high comodulation between different frequency ranges, and to have a comparatively long decay time of the post masking modeling in case of a comparatively low comodulation between different frequency ranges, which has been found to well reflect the behavior of the human auditory system. This is due to the recognition that this concept is well suited to reflect or model the so-called dip listening effect of the human auditory system.

[0028] Worded yet differently, in this embodiment, the decay time of the post masking modeling may be adjusted in such a manner that it is well adapted to the human hearing, since the concept takes into account a variability of the decay time of the post masking modeling and the dependency of the decay time of the post masking modeling from the comodulation. As a consequence, the resulting masking threshold information, which is obtained using the application of the post masking modeling with adjustable decay time is particularly well adapted to an actual masking threshold occurring in a human auditory system.

[0029] Furthermore, it has been recognized that both the derivation of the decay time of the post masking modeling and the realization of a post masking modeling using the variable decay time can be implemented with comparatively low complexity, e.g. since the decay time may be derived from a value (e.g. from a single scalar value) representing a comodulation strength using a scalar (and possibly memory-less) mapping function of relatively low complexity, e.g. using a relatively simple linear or piecewise-linear mapping function. Furthermore, it has also been recognized that the application of the post masking modeling with variable decay time can also be implemented with low computational complexity, e.g.

[0030] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX oiMiliE

[0031] 5

[0032] using a simple recursive mapping approach which only becomes effective for a reduction (or significant reduction) of an (e.g. preliminary) masking threshold value which is input into the post-masking modeling.

[0033] To conclude, the embodiment described is well suited to provide masking threshold values of high quality while keeping a computational complexity reasonably small.

[0034] In accordance with embodiments of the present inventive concept, the masking threshold determinator is configured to scale (e.g., attenuate) (e.g., in a per-frequency-range-wise manner) the envelope information (e.g., ek(m)), or a pre-processed version thereof (e.g., Sfc(m)), or the post-masking processed envelope information (e.g., pfc(m)), in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mk(m) or in dependence on a comodulation), e.g., using an attenuation factor fl'k(m). For example, this (e.g. this concept) may be a variant, in which the CMR estimation disclosed herein (e.g. the different concepts for the CMR estimation disclosed herein) can be used (e.g. in published methods, e.g. in published methods according to [3] and / or [4]) for a direct reduction of the masking thresholds.

[0035] This embodiment of the invention is based on the finding, that by additionally scaling the envelope information, the preprocessed version thereof or the post-masking processed envelope information can be a further advantageous utilization of the comodulation strength information to achieve a higher quality masking threshold. With this scaling concept, a better adaption of the masking threshold to the human earing can be achieved. Since the comodulation strength information contains information of the comodulation between a current frequency band and other frequency bands, the masking threshold determinator can use the comodulation strength information to adapt the masking threshold of a current frequency band, such that sounds, that have comodulated properties can be better processed or encoded using this embodiment of the invention.

[0036] Since the comodulation strength information is already present and does not need to be calculated differently, the scaling of the envelope information, a preprocessed version thereof or the post masking processed envelope information requires minimal additional computational resources and therefore depicts an improved compromise between a quality of a determination of a masking threshold, an implementation complexity and possibly an achievable perceptual audio quality.

[0037] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0038] O IlliiliS

[0039] 6

[0040] Moreover, it has been recognized that by combining the comodulation-dependent adjustment of the decay time of the post-masking modeling and the comodulationdependent scaling of the envelope information, good dynamic and static (or quasi-static) characteristics of the masking threshold information can be achieved with moderate computational effort. In particular, it has been recognized that the combination of the scaling of the envelope information and of the adjustment of the decay time of the post-masking modeling does not interfere but rather provides a particularly good result.

[0041] Such a combination of the comodulation-dependent adjustment of the decay time and of the post-masking modeling and the comodulation-dependent scaling of the envelope information may, for example, correspond to (or may, for example, be achieved by) a combination of the processing of Figures 4 and 5. For example, Figure 4 is a variant in which the CMR estimation disclosed herein may be used in already published methods (see, for example, [3] and [4]) for a direct reduction of the masking thresholds. For example, Figure 5 is a new concept and disclosed herein in detail.

[0042] According to embodiments the masking threshold determinator is configured to scale (e.g., attenuate), e.g., in a per-frequency-range-wise manner, the envelope information (e.g., efc(m)), or the pre-processed version thereof (e.g., ek(m)), or the post-masking processed envelope information (e.g., pfc(m)), in dependence on the comodulation strength information using an attenuation factor (e.g., using an attenuation factor gk(m),- e.g., using an atenuation factor which serves as a scaling factor; e.g., using an attenuation factor which is smaller than or equal to 1 ),

[0043] wherein the masking threshold determinator is configured to derive the attenuation factor (e.g., 9km)onthe basis of a comodulation coefficient (e.g., cfe(n)), for example, on the basis of a comodulation coefficient associated with a currently considered frequency range, e.g., solely on the basis of the comodulation coefficient associated with the currently considered frequency range, for example, using a linear mapping, e.g., using a linear mapping which maps a comodulation coefficient associated with a currently considered frequency range onto an attenuation factor value associated with the currently considered frequency range, e.g., using a linear mapping which maps a maximum value of the comodulation coefficient associated with a currently considered frequency range onto a minimum value of the atenuation factor, and which maps a minimum value of the comodulation coefficient associated with the currently considered frequency range onto a maximum value of the attenuation factor, or for example, using a non-linear mapping, e.g.,

[0044] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX ga^wiai

[0045] 7

[0046] using a non-linear mapping which maps a comodulation coefficient associated with a currently considered frequency range onto an atenuation factor value associated with the currently considered frequency range, e.g., using a non-linear mapping which maps a maximum value of the comodulation coefficient associated with a currently considered frequency range onto a minimum value of the attenuation factor, and which maps a minimum value of the comodulation coefficient associated with the currently considered frequency range onto a maximum value of the attenuation factor.

[0047] This embodiment of the invention is based on the finding, that additionally scaling the envelope information, the preprocessed version thereof or the post-masking processed envelope information using an attenuation factor, that was obtained on the basis of a comodulation coefficient, can be a further advantageous utilization of the comodulation strength information to achieve a higher quality masking threshold.

[0048] By attenuating one or more of these information items, a better adaption to the human earing can be achieved. Since the comodulation strength information contains the information of correlation between a current frequency band and other frequency bands, the masking threshold determinator can, for example, use this comodulation strength information to lower the masking threshold for sounds that have comodulated properties. This behavior is inspired by the human auditory system, which can use comodulated frequency bands (e.g., bands that fluctuate in amplitude together) to detect a target signal (e.g., a signal of the current frequency band) in the presence of masking noise.

[0049] Furthermore, the usage of an attenuation factor as described in this embodiment of the invention represents a highly efficient form of implementation, since it may, for example, be represented as single scalar value derived from the comodulation strength information. The comodulation strength information is already being calculated and therefore, the attenuation of the envelope information, a preprocessed version thereof or the post masking processed envelope information requires minimal additional computational resources.

[0050] Moreover, it has been recognized that by combining the comodulation-dependent adjustment of the decay time of the post-masking modeling and the comodulation¬ dependent scaling of the envelope information, good dynamic and static (or quasi-static) characteristics of the masking threshold information can be achieved with moderate computational effort. In particular, it has been recognized that the combination of the scaling of the envelope information and of the adjustment of the decay time of the post-masking modeling does not interfere but rather provides a particularly good result.

[0051] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX . S

[0052] 8

[0053] Such a combination of the comodulation-dependent adjustment of the decay time and of the post-masking modeling and the comodulation-dependent scaling of the envelope information may, for example, correspond to (or may, for example, be achieved by) a combination of the processing of Figures 4 and 5. For example, Figure 4 is a variant in which the CMR estimation disclosed herein may be used in already published methods (see, for example, [3] and [4]) for a direct reduction of the masking thresholds. For example, Figure 5 is a new concept and disclosed herein in detail.

[0054] In accordance with embodiments of the present inventive concept, the post-masking modeling is configured to perform an envelope following with an adjustable decay time, for example, wherein the adjustable decay time may be equal to the post-masking decay time. This embodiment of the invention is based on the finding, that the perceptual audio quality increases, if the post-masking modeling is configured such that it has little to no delay in the case of an increase of the envelope, while its decay time can be adjusted (and therefore, for example, may bring along an adjustable slow-down of changes of the masking threshold). Since no additional information needs to be computed for this embodiment of the invention, it can be easily implemented in a masking threshold determinator. Furthermore, since it has been recognized that the processes as described in this embodiment are adapted to the natural human hearing system, a masking threshold determinator according to this embodiment can lead to reliable masking threshold information and also to good achievable perceptual audio quality.

[0055] In accordance with embodiments of the present inventive concept, the post-masking modeling is configured such that an output signal of the post-masking modeling follows an increase of an input signal of the post-masking modeling without delay, for example, without an (intentionally) added delay that goes beyond an imminent delay caused by a timediscrete implementation and / or caused by a limited processing speed; e.g., in a time discrete manner; e.g., according to pk(m) = ek' (m), and

[0056] such that the output signal of the post-masking modeling follows a reduction of the input signal of the post-masking modeling with a (limited) maximum step size which is an adjustable fraction (e.g., (1 - dfc)pfc(m - 1)) of a of a previous output signal value (e.g., pk(m - 1)) of the post-masking modeling determined by a decay factor (e.g., dfe),

[0057] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX aMMIIIIC.

[0058] 9

[0059] wherein the masking threshold determinator is configured to adjust the decay factor in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mfcor in dependence on a comodulation coefficient ck' ) (e.g., to thereby adapt the post-masking decay time of the post-masking modeling).

[0060] This embodiment of the invention is based on the finding, that an excessively fast decay of the output signal can lead to inappropriate masking threshold information and consequently to perceptual bad audio quality, while for the increase of an input signal an immediate increase of the output signal of the post-masking modeling should be ensured to obtain a reliable masking threshold information and consequently to ensure high perceptual audio quality. Furthermore, it has been found, that selecting the decay factor in dependence on the comodulation strength information can be further advantageous for the quality of the resulting masking threshold.

[0061] For example, the concept allows to adjust (and reduce) the maximal speed of reduction of the output signal of the post-masking modeling, with which the characteristics of the human auditory system can be exploited. It has been recognized that the human auditory system is not well adapted to vary fast reductions of an audio signal, and that the reaction time of the human auditory system to reductions of a signal level is dependent on comodulation characteristics. The masking threshold determinator can therefore adapt the temporal evolution of the masking thresholds for comodulated frequency bands of the audio signal. This can, for instance, result in a coarser quantization of an audio signal, where the human auditory system is unable to resolve the signal with the same level of detail as it can at other times.

[0062] Furthermore, this concept is associated very little computational cost, since all the information is already present and as for calculations only simple operations need to be performed.

[0063] To conclude, this embodiment allows with minimal computational cost the adaption of a masking threshold to beter depict and exploit the inner workings of the human auditory system, leading to higher quality masking thresholds.

[0064] According to embodiments of the invention, the post-masking modeling is configured such that an output signal of the post-masking modeling follows an increase of an input signal of the post-masking modeling without delay (e.g., without an (intentionally) added delay that goes beyond an imminent delay caused by a time-discrete implementation and / or caused

[0065] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX SBIIBI1S.! «.

[0066] 10

[0067] by a limited processing speed; e.g., in a time discrete manner; e.g., according to pk(m) = e (m)), and

[0068] such that the output signal of the post-masking modeling follows a step-wise reduction of the input signal of the post-masking modeling with an exponential decay (e.g., with a step- 5 wise exponential decay; e.g., with a decay of the form pk(m) = dkpk(m - 1)), wherein the masking threshold determinator is configured to adjust a time constant of the exponential decay (e.g., a decay factor dk) in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mkor in dependence on a comodulation coefficient ck) (e.g., to thereby adapt the post-masking decay time of the 10 post-masking modeling).

[0069] This embodiment of the invention is based on the finding, that based on the circumstances, an exponential decay of the post masking modelling can be advantageous compared to a linear decay of an output signal. This concept exploits the fact, that the human auditory system does not work on a completely linear level. While linear implementations are more 15 cost efficient than exponential ones, the increase in quality of the masking threshold determinator is worth the little increase in computational cost. Thus, the embodiment described provides an improved compromise between the quality of a determination of a masking threshold and computational complexity required for its implementation.

[0070] 20 According to embodiments of the invention, the post-masking modeling is configured to obtain the output signal (e.g., pk(m)), for example, a current output signal, associated with a current time index value, of the post-masking modeling using a maximum function which determines a maximum value out of a current envelope information (e.g., ek(n)), or a pre- processed version thereof (e.g., e (n)), and a scaled version of a previous output signal 25 (e.g., associated with a previous time index value) of the post-masking modeling (e.g., dkpk(m - 1)),

[0071] wherein the post masking modeling is configured to scale a previous output signal of the post-masking modeling (e.g., pk(m - 1)) using a decay factor (e.g., dk(m)) (e.g., a scaling value or a time-variant scaling value) which is determined in dependence on the 30 comodulation strength information (wherein the decay factor determines the post-masking decay time), to obtain the scaled version (e.g., dkpk(m - 1)) of the previous output signal of the post-masking modeling.

[0072] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX ;!■■■&. a i

[0073] 11

[0074] This embodiment of the invention is based on the finding, that utilizing a maximum value of a current envelope information or a pre-processed version thereof and of a scaled version of a previous output signal of the post-masking modeling can be a computationally efficient way of determining a masking threshold with a high quality,

[0075] 5

[0076] The application of a maximum operator poses no significant computational cost to a masking threshold determinator. Furthermore, the usage of the maximum of, for example, the current envelope information, or a preprocessed version thereof, and a scaled version of a previous output signal can lead to a higher masking threshold in situations, where the 10 envelope information or a preprocessed version thereof are already faded,, but a higher masking threshold can still be advantageous in terms of the bitrate required to achieve a desired perceptual audio quality. For example, even though using the masking threshold values obtained using this concept may not directly improve an audio quality (e.g. of an audio encoding / decoding performed using the masking threshold), a bitrate required to 15 obtain a certain quality (e.g. a certain quality level) may be improved (e.g. reduced) by using the obtained masking threshold values (e.g. in an audio encoder).

[0077] Thus, this embodiment can lead to an improvement of the compromise between computational complexity, a quality of a determination of a masking threshold, and an achievable perceptual audio quality.

[0078] 20

[0079] In accordance with embodiments of the present inventive concept, the post-masking modeling is configured to obtain the output signal pk(m) of the post-masking modeling according to pk(m) = max(e (m,dkpk(m - 1)) wherein e (m) is an envelope information (e.g., a value describing an envelope in a frequency range having frequency range index k 25 for time index m), or a pre-processed (e.g., averaged and / or low-pass-filtered and / or downsampled) version thereof, wherein pk(m - 1) is a previous value of the output signal of the post-masking modeling, wherein dkis a decay factor which is adjusted in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mkor in dependence on a comodulation coefficient c ) (e.g., to thereby adapt the 30 post-masking decay time of the post-masking modeling); wherein k is a frequency index, for example, a frequency range index designating a frequency range or a frequency band index designating a frequency band;

[0080] wherein m is a time index; and wherein max(...) is a maximum value operator.

[0081] )

[0082] !

[0083] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX This embodiment of the invention is based on the finding, that utilizing a maximum value of a current envelope information or a pre-processed version thereof and of a scaled version of a previous output signal of the post-masking modeling can be a computationally efficient way of determining a masking threshold with high quality.

[0084] The application of this explicit maximum operator poses no significant computational cost to a masking threshold determinator and constitutes an especially efficient implementation. Furthermore, the usage of the maximum of, for example, the current envelope information, or the preprocessed version thereof, and of a scaled version of a previous output signal can lead to a higher masking threshold in situations, where the envelope information or a preprocessed version thereof are already faded, but a higher masking threshold can still be advantageous in terms of the bitrate required to achieve a desired perceptual audio quality. Thus, this embodiment can lead to an improvement of the compromise between computational complexity, a quality of a determination of a masking threshold, and an achievable perceptual audio quality.

[0085] According to embodiments the masking threshold determinator is configured to obtain a plurality of sequences of frequency band values representing an audio content in different frequency ranges (e.g., using a filter bank; e.g., using a filterbank which models a spectral decomposition in an inner ear; e.g., using a plurality of bandpass filters; e.g., using a plurality of bandpass filters with center frequencies uniformly distributed on a perceptual frequency scale like ERB or BARK; e.g., using a plurality of Gammatone filters); and

[0086] wherein the masking threshold determinator is configured to derive the envelope information (e.g., ek(n) or e (n)) from the sequences of frequency band values (e.g., using a magnitude determination; e.g., using a rectification and a smoothening in the case of real- valued frequency band values; e.g., using a determination of magnitudes (of complex valued frequency band values) and using a smoothing of said magnitudes in the case of complex-valued frequency band values).

[0087] This embodiment of the invention is based on the finding, that combining information of a plurality of frequency band values can be advantageous for the quality of a determination of a masking threshold. For example, this concept allows the usage of a filterbank to decompose an audio signal, and this filterbank can be designed specifically to model the spectral decomposition in a human inner ear. This application of knowledge of the human

[0088] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX saiBBiBr..

[0089] auditory system in return ailows for the design of a higher quality of determination of a masking threshold, since underlying knowledge about the human biology is used for its design.

[0090] 5 According to embodiments of the invention, the masking threshold determinator is configured to obtain the comodulation strength information using an evaluation of a covariance (e.g., a cross-covariance) or of a correlation (e.g., a cross-correlation) between temporal envelopes (e.g., normalized temporal envelopes) in (e.g., associated with) two or more different frequency ranges (e.g., in neighboring frequency ranges or in distant 0 frequency ranges, e.g., with one or more other frequency ranges in between), e.g., using an evaluation of a plurality of cross-covariances or of cross-correlations between respective pairs of temporal envelopes having frequency range indices (e.g., band indices) comprising a predetermined difference (e.g., d) which is larger than 1.

[0091] This embodiment of the invention is based on the findings, that utilizing different frequency 5 ranges to obtain the comodulation strength information can be advantageous for quality of determination of a masking threshold. By utilizing different frequency ranges to obtain the comodulation strength information, the comodulation strength information can be determined in an accurate way, since the information it relies on is more expansive. Also, determining a covariance or a correlation can be performed in a computationally efficient 0 manner, both in hardware and in software depending on the requirements.

[0092] Furthermore, this implementation poses little additional computational cost to the masking threshold determinator, since, for example, the different frequency ranges that are considered are already being processed to determine masking thresholds for said frequency bands. Therefore, the available information can be utilized with no significant 5 computational cost to obtain an accurate comodulation strength information, which in return can lead to a better determination of a masking threshold.

[0093] This typically results in good quality masking threshold values and high computational efficiency.

[0094] 0 According to embodiments of the invention, the masking threshold determinator may be configured to determine the comodulation strength information (e.g., a plurality of comodulation strength values mkor a plurality of comodulation coefficients c'k) using a computation of a correlation coefficient Ck'laccording to

[0095] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX C / cz (n)

[0096] Ckt =

[0097] ek(n) eE2(n))

[0098]

[0099] wherein

[0100] C

[0101]

[0102] kl(n) = e^(nMn) - ak(n)al(n')

[0103] wherein

[0104] afc(n) = e^n)

[0105] wherein

[0106] az(n) = e nj

[0107] wherein k is a frequency range index,

[0108] wherein I is a frequency range index, with I = k - d or I = k + d,

[0109] wherein d is a predetermined distance value,

[0110] wherein n is a time index,

[0111] wherein ek(n) is an envelope value in a frequency band having a frequency band index fc; wherein et(n) is an envelope value in a frequency band having a frequency band index l wherein (...) is an averaging operation which performs a temporal averaging (e.g., a moving average averaging, ora UR lowpass averaging) (e.g., overtime) (e.g., of a sequence efc(n) of envelope values in the frequency band having frequency band index k, of a sequence eE(n) of envelope values in the frequency band having frequency band index I, or of a sequence of values resulting from an element-wise multiplication of the sequence ekQ and of the sequence e^n) or of a sequence e (n) or of a sequence e (n)).

[0112] This embodiment of the invention has been found to be especially advantageous, since it determines the comodulation strength information using a computation of a correlation coefficient, leading to an accurate determination of a masking threshold. Also, an appropriate normalization is included in the above computation, which improves independence of the results from a loudness of the signals. All formulas described above can be implemented in a masking threshold determinator with little computational cost

[0113] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX associated with them, therefore this embodiment of the invention depicts an improved compromise between computational complexity and a high-quality determination of a masking threshold.

[0114] According to embodiments of the invention, the masking threshold determinator may be configured to determine the comodulation strength information (e.g., a plurality of comodulation strength values mkor a plurality of comodulation coefficients c ) using a computation of a comodulation strength value mkaccording to

[0115] z \ 4- VQcuCn))

[0116] m

[0117]

[0118] kW~ 2ak(n) + afn) + au(n)

[0119] or according to

[0120] G (-A / Cfc / Cn) _|_

[0121] (n) = — — — — — — — — • — — i—

[0122]

[0123] max(ak(n), ai{n),au(n))

[0124] (here G may be a normalization factor, e.g., a predetermined value. For example, G may be equal to 2.8 in the first alternative of the mkcalculation, and G may, for example, be equal to 0.7 in the second alternative of the mkcalculation)

[0125] or according to

[0126] m2^Ckjfn)

[0127] (jQ = - if k < d

[0128] with j =

[0129]

[0130] afc(n) + aj(n) if k > K — d or according to

[0131] ,, yfcj^j (n ) (u if k < d mk(,n) = with] =,

[0132]

[0133] max(ak{n),aj(n)) U if k > K — d

[0134] wherein

[0135] Cki(n) = ek(n e((n) - ak(n) at(n)

[0136] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX wherein Ckuand CkJ- are defined like Ckl, wherein frequency index u or frequency index j takes the place of frequency index I,

[0137] wherein

[0138] afe(n) = efc(n)

[0139] wherein

[0140] a;(n) = et(n)

[0141] wherein auand ctj are defined like at, wherein frequency index u or frequency index j takes the place of frequency index I

[0142] wherein k is a frequency range index,

[0143] wherein I is a frequency range index, with k = k - dt,

[0144] wherein u is a frequency range index, with u = k + d2,

[0145] wherein dzand d2are predetermined distance values (wherein dtand d2may be equal, e.g., equal to d),

[0146] wherein n is a time index,

[0147] wherein ek(n) is an envelope value in a frequency band having a frequency band index fc; wherein et(n) is an envelope value in a frequency band having a frequency band index wherein eu(n) is an envelope value in a frequency band having a frequency band index u; wherein (77) is an averaging operation which performs a temporal averaging (e.g., a moving average averaging, or a HR lowpass averaging) (e.g., overtime) (e.g., of a sequence ek(n) of envelope values in the frequency band having frequency band index k, of a sequence e|(n) of envelope values in the frequency band having frequency band index I, of a sequence eu(n) of envelope values in the frequency band having frequency band index u, or of a sequence of values resulting from an element-wise multiplication of the sequence ekQ and of the sequence e^n), or of a sequence of values resulting from an element-wise multiplication of the sequence ek() and of the sequence eu(n)); herein ekQ may also refer to efc(n).

[0148] wherein G is a normalization factor (e.g., a predetermined value).

[0149] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0150] (4 This embodiment of the invention can be especially advantageous for determination of a masking threshold since it has been found, that the defined calculations strike an optimal compromise between computational complexity and quality of determination of a masking threshold. The defined calculations rely on an upper and lower neighboring frequency band, with a defined distance to the current frequency band, taking into account three different frequency bands, which is sufficient enough information for an accurate determination of a masking threshold and is still comparatively cheap in terms of associated computational cost.

[0151] Due to the accounting of one upper and one lower frequency band, the masking threshold determinator according to this embodiment of the invention prevents the emergence of bias errors.

[0152] Furthermore, this embodiment additionally considers the edge cases, where the upper or lower neighbor of a defined distance does not lie in the available sequence of frequency bands. Thus, the described embodiment of the invention can be utilized across all frequency bands and exhibits no shortcomings in this respect.

[0153] Since all operations in the described embodiment are based on fairly simple mathematical operations, the computational complexity remains low and its implementation does not pose significant computational difficulties.

[0154] The defined alternative calculations of mk, once for two possible neighbors and once for the edge case, utilize a maximum operator, since it has been found, that this can result in a favorable comodulation strength information mk.

[0155] To conclude, the embodiment described is well suited to provide masking threshold values of high quality while keeping a computational complexity reasonably small.

[0156] According to embodiments of the invention, the masking threshold determinator is configured to determine the comodulation strength information (e.g., a plurality of comodulation strength values mkor a plurality of comodulation coefficients c ) using a computation of a comodulation strength value,

[0157] wherein the computation of the comodulation strength value (e.g., mfc(n)) comprises a computation of a quotient between

[0158] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX - a sum of a plurality of covariance values describing a respective covariance (e.g., a cross¬ covariance) or a respective correlation (e.g., a cross-correlation) between temporal envelopes (e.g., normalized temporal envelopes) (e.g., represented by respective sequences of envelope values) in two different frequency ranges (e.g., in neighboring frequency ranges or in distant frequency ranges, e.g., with one or more other frequency ranges in between), or a sum of a plurality of exponentiated versions (e.g., exponentiated with exponents smaller than 1, or exponentiated with exponents of 1 / 2) of the covariance values (e.g., a sum of square roots of the covariance values CMand Cku), and

[0159] - a sum, or a weighted sum (e.g., 2ak(n) + a;(n) + au(n)), or a maximum (e.g. a maximum of ak, ai and au, e.g. max(^akn), au(n), ai(n)) of average values (e.g., ak, ai, au) of sequences of envelope values (e.g., ek, ebeM) in the different frequency ranges (e.g., having frequency range indices k, I and u).

[0160] This embodiment of the invention can lead to a higher quality of determination of a masking threshold since it has been found, that the adaption of the calculation of the comodulation strength information depending on the current circumstances can be highly advantageous. Since this embodiment allows two ways of calculating the comodulation strength value, a masking threshold determinator can be configured to utilize the method of computation, which suits the current circumstances the best. This enables the creation of a masking threshold determinator which achieves an optimal compromise between a quality of a determination of a masking threshold, an implementation complexity and possibly an achievable perceptual audio quality.

[0161] According to embodiments of the invention, the masking threshold determinator is configured to map a comodulation strength value (e.g., mfc(n)) onto a comodulation coefficient (e.g., ck(n)) using a linear mapping and a limitation of a range of values (e.g., such that the comodulation coefficient is limited to lie in a predetermined range of values, e.g., in a predetermined closed interval, e.g., in closed interval between 0 and 1 (including 0 and 1)).

[0162] This embodiment of the invention is based on the finding, that using a linear mapping with a limitation of a range of values to map a comodulation strength value onto a comodulation coefficient is a highly efficient way of determining an accurate comodulation coefficient. Since such an approach represents minimal computational cost and since it has been found, that the resulting comodulation coefficients can lead to a high quality of masking

[0163] FH1IS24EM24 FH1IS24EM24-2024346428. DOCX thresholds, this embodiment represents an improved compromise between a quality of a determination of a masking threshold, an implementation complexity and possibly an achievable perceptual audio quality.

[0164] In accordance with embodiments of the present inventive concept, the masking threshold determinator is configured to use a smaller rate of samples in the post-masking modelling than in the determination of the comodulation strength information, it is to be noted, that the post-masking modelling may, for example, also be exchanged with a post-masking evaluation.

[0165] This embodiment of the invention is based on the finding, that a smaller rate of samples in the post-masking evaluation than in the determination of the comodulation strength information can be used, while maintaining a high level of achievable perceptual audio quality. Therefore, a number of computations can be reduced, while keeping a high level of achievable perceptual audio quality, constituting an improved compromise between computational cost and achievable perceptual audio quality.

[0166] In accordance with embodiments of the present inventive concept, the masking threshold determinator is configured to apply an averaging and a downsampling to the envelope information (e.g., efc(n)), in order to obtain a pre-processed version of the envelope information (e.g., having a smaller temporal resolution than the envelope information) (e.g., using a sample rate which is smaller than an original sample rate of the envelope information),

[0167] wherein the masking threshold determinator is configured to apply the post-masking modeling to the pre-processed version of the envelope information, to obtain the postmasking-processed envelope information; and

[0168] wherein the masking threshold determinator is configured to determine the comodulation strength information on the basis of the envelope information, using a full temporal resolution of the envelope information (e.g., using a sample rate which is identical to an original sample rate of the envelope information).

[0169] This embodiment of the invention is based on the finding, that this utilization of averaging and downsampling for processing the envelope information while keeping the full temporal resolution for the determination of the comodulation strength information can lead to an improved compromise between perceptual audio quality and computational complexity.

[0170] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 20

[0171] Leaving out information wherever possible, while keeping the full information for all relevant computations depicts an important, non-trivial step in this embodiment.

[0172] According to embodiments of the present invention, the masking threshold determinator is configured to obtain a decay time constant value (e.g., varDecayTimeConst) (which may, for example, determine a post-masking decay time of the post-masking modeling) in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mk, or in dependence on a comodulation coefficients ck' ) using a linear mapping.

[0173] This embodiment of the invention is based on the finding, that by utilizing a linear mapping for obtaining a decay time constant value in dependence on the comodulation strength information complex computations can be evaded, while still being accurate enough to impose no loss of perceptual audio quality. Therefore, an improved compromise between computational complexity and perceptual audio quality can be achieved.

[0174] According to embodiments of the present invention, the masking threshold determinator is configured to obtain a decay time constant value, which determines a post-masking decay time of the post-masking modeling, in dependence on a comodulation coefficient (e.g., ck) using a linear mapping

[0175] (wherein, for example, the masking threshold determinator is configured to linearly map a range of values of the comodulation coefficient between 0 and 1 onto a range of values of the decay time constant value between a predetermined minimum decay time constant and a predetermined maximum decay time constant;

[0176] wherein, for example, a comodulation coefficient of 0 is mapped onto the predetermined maximum decay time constant, and wherein, for example, a comodulation coefficient of 1 is mapped onto the predetermined minimum decay time coefficient;

[0177] wherein, for example, the predetermined decay time coefficient is in a range between 1ms and 5ms, and wherein, for example, the predetermined maximum decay time constant is in a range between 7ms and 20ms, or even in a range between 7ms and 60ms). For example, it has been recognized that, in some cases, higher values for the maximum decay time (e.g. higher than 20ms) may bring along good results.

[0178] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX This embodiment of the invention is based on the finding, that the usage of a linear mapping to obtain a decay time constant value, which determines a post-masking decay time of the post-masking modeling in dependence on a comodulation coefficient is advantageous, since it poses minimal computational complexity while it can lead to high perceptual audio quality. In specific embodiments of this invention, the linear map can be utilized to restrict the mapping onto a defined limited range which can lead to a precise behaviour of the masking threshold determinator, allowing for an optimal adjustment of its parameters to provide the best perceptual audio quality.

[0179] According to embodiments of the present invention, the masking threshold determinator is configured to scale a previous output signal of the post-masking modeling (e.g., pk(m - 1)) using a decay factor (e.g., dk(m)) (e.g., a scaling value or a time-variant scaling value) (wherein the decay factor determines the post-masking decay time), to obtain a scaled version (e.g., dkpk(m - 1)) of the previous output signal of the post-masking modeling, and to wherein the masking threshold determinator is configured to use the scaled version of the previous output signal of the post-masking modeling as a current output signal of the post-masking modeling when slowing down a decay of the output signal of the post-masking modeling (e.g., in case of a fast decay of an input signal of the post-masking modeling), for example in order to control a post masking decay time,

[0180] wherein the masking threshold determinator is configured to determine the decay factor in dependence on a decay time constant value (e.g., in dependence on the decay time constant value mentioned before).

[0181] This embodiment of the invention is based on the finding, that a version of a previous output signal can be scaled and utilized as a current output signal of the post-masking modeling- Due to the dependence of the decay factor on a decay time constant one can ensure, that excessive drops of the output signal, which can lead to unreliable masking threshold values and consequently to bad perceptual audio quality, can be prevented and the masking threshold determinator can be configured, such that its masking thresholds fit nicely to the human auditory system. Additionally, by controlling the decay through a (e.g single) scalar value, low computational complexity can be ensured. Furthermore, this embodiment reuses already existing information and therefore requires little computation. In conclusion, this embodiment represents an improved compromise between perceptual audio quality and computational complexity.

[0182] FHHS24EM24 FHIIS24EM24-2024346428. DOCX 22

[0183] According to embodiments of the present invention, the masking threshold determinator is configured to determine the decay factor (e.g., dfc) in dependence on the decay time constant value, such that the decay factor increases with increasing decay time constant value.

[0184] This embodiment of the invention is based on the finding, that this connection of the decay factor with the decay time constant value can be advantageous for obtaining reliable masking threshold values and consequently for achieving a good perceptual audio quality. By increasing the decay factor with an increasing decay time constant value, the logically most plausible connection is formed between those two values.

[0185] According to embodiments of the present invention, the masking threshold determinator is configured obtain the decay factor (e.g., dfe) using a subtraction of a value (e.g., — — — — — • ) representing a quotient between a sample period value (e.g., varDecayTimeConst * sampleRate

[0186] - 1— or — ) and a decay time constant value (e.g., varDecayTimeConst) from sampleRate sampleRateJ

[0187] a predetermined value (e.g.1) (e.g., such that a ratio between a sample period time (duration) and a decay time constant is subtracted from a predetermined value, e.g., 1, to obtain the decay factor).

[0188] This embodiment of the invention is based on the findings, that calculating the decay factor by subtracting a quotient between a sample period value and a decay time constant value from a predetermined value can be advantageous to efficiently derive meaningful masking threshold values that allow for obtaining a good perceptual audio quality.

[0189] According to embodiments of the present invention, the masking threshold determinator is configured to obtain the decay factor dkaccording to

[0190] d -1000

[0191] varDecayTimeConst*sampleRate'

[0192] wherein varDecayTimeConst is a decay time constant value (e.g., describing the desired decay time in milliseconds),

[0193] wherein sampleRate is a value describing a sample rate of the output signal of the post- masking modeling (e.g., describing the sample rate in samples per second).

[0194] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX This embodiment of the invention is based on the finding, that calculating the decay factor by subtracting quotient between a sample period value and a decay time constant value from 1 can be advantageous to efficiently obtain a decay factor d which brings along good results for the determination of the masking threshold values. Furthermore, basing the Variable sampleRate on the sampling performed at the output of the post-masking modeling can be advantageous, since it imposes minimal computational cost to the masking threshold determinator.

[0195] Additionally, it is to be noted, that, alternatively, the sampleRate may be derived from the sampling rate of the input signal.

[0196] When using the input sample rate for the calculation of dkand in the case of downsampling, dkmay need to be further processed, for example, using dkto the power of a downsampling factor, wherein the downsampling factor is a factor, describing the downsampling of the signal.

[0197] In accordance with an aspect of the present inventive concept, an embodiment creates a quantization step size determinator for providing a quantization step size information (e.g., qi, qj, also designated as qt, q,) on the basis of an input audio signal,

[0198] wherein the quantization step size determinator comprises a masking threshold determinator according to any of the previously mentioned embodiments of this invention, wherein the masking threshold determinator is configured to determine a masking threshold information (e.g., pkrn)) on the basis of an input audio signal (e.g., x(n)),

[0199] wherein the quantization step size determinator is configured to determine the quantization step size information using the masking threshold information (e.g., using a post-masking-processed envelope information).

[0200] This embodiment of the invention is based on the findings, that the high quality masking threshold information, obtained by a masking threshold determinator according to an embodiment of this invention, can be used by a quantization step size determinator to achieve good quantization step sizes. This in return can lead to a good achievable perceptual audio quality at certain bit rates.

[0201] With the improved masking threshold, the quantization step size determinator can determine the quantization step sizes more precisely, and can determine the step sizes such that the quantization resolution is high, where the human auditory system can utilize the high resolution and low, when the human auditory system is not capable of utilizing a

[0202] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 24

[0203] encoding a signal using a quantization step size determined by the quantization step size determinator according to embodiments of this invention.

[0204] To conclude, the improved step sizes determined by a quantization step size determinator according to an embodiment of this invention can be used to achieve a higher perceptual 5 audio quality, and it can be attributed to the higher quality masking thresholds obtained by a masking threshold determinator according to this invention.

[0205] According to embodiments of the present invention, the quantization step size determinator is configured to determine the masking threshold information using a first frequency 10 resolution, and

[0206] wherein the quantization step size determinator is configured to determine the quantization step size information using a second frequency resolution, which is different from the first frequency resolution.

[0207] This embodiment of the invention is based on the finding, that by utilizing a first and second 15 frequency resolution that are different from each other, a quantization step size determinator can use appropriate frequency ranges to determine the optimal quantization step sizes in terms of achievable perceptual audio quality. This can be attributed to the fact, that knowing how the audio signal is transformed into different frequency resolutions can be important information to determine the quantization step size. Also, the masking threshold values can 20 be determined using a frequency resolution which is well adapted to an impact of a comodulation on the human hearing. On the other side, the quantization threshold values can be determined using a frequency resolution which is adapted to the characteristics of the audio coding, e.g., to widths of scale factor bands. For example, a mapping can made from the masking threshold values onto the quantization step size values, taking into 25 account the different frequency resolutions. Thus, the quantization step sizes can be obtained both efficiently and accurately.

[0208] An embodiment according to the present invention creates an audio encoder for providing an encoded audio representation on the basis of one or more input audio signals,

[0209] 30 wherein the audio encoder comprises a quantization step size determinator according to any of the previously mentioned quantization step size determinator, wherein the quantization step size determinator is configured to obtain the quantization step size information on the basis of the one or more input audio signals;

[0210] FHIIS24EM24 FHHS24EM24-2024346428. DOCX wherein the audio encoder comprises a quantizer configured to quantize different frequency ranges of the one or more input audio signals (e.g., spectral domain coefficients (e.g., MDCT coefficients) representing the different frequency ranges, or sub-band signals representing the different frequency ranges) using different effective quantization resolutions (e.g., using a frequency-dependent scaling, to obtain a scaled representation of the different frequency ranges of the one or more input audio signals, and using a quantization of the scaled representation of the different frequency ranges, wherein the frequency-dependent scaling is controlled in dependence on the quantization step size information provided by the quantization step size determinator), to obtain quantized representations of the different frequency ranges of the one or more input audio signals; and

[0211] wherein the audio encoder is configured to encode (e.g., to losslessly encode) the quantized representations of the different frequency ranges of the one or more input audio signals, to obtain the encoded representation of the one or more input audio signals.

[0212] This embodiment of the invention is based on the finding, that by utilizing a quantization step size determinator according to embodiments of this invention, it is possible to achieve a good perceptual audio quality of an audio signal encoded using an encoder as described in this embodiment of the invention. This can be attributed to the high quality step sizes, obtained by a quantization step size determinator according to embodiments of this invention, which in return can be attributed to the high quality masking thresholds, obtained with a masking threshold determinator according to embodiments of this invention.

[0213] The high quality masking thresholds can be utilized by the quantization step size determinator to determine the step sizes such that the resolution is always optimized for the human auditory system. This can lead to a high achievable perceptual audio quality when working at certain bit rates.

[0214] An embodiment according to the present invention creates a method for providing a masking threshold information (e.g., pfc(m)) on the basis of an input audio signal (e.g., x(n)), wherein the method comprises obtaining envelope information (e.g., a plurality of envelope signals; e.g., efe(n); e.g., respective envelope signals associated with different frequency ranges) describing envelopes of different (e.g., overlapping or non-overlapping) frequency ranges (e.g., different frequency bands) of the input audio signal (e.g x(n)).

[0215] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0216] C 1 The method comprises determining a comodulation strength information (e.g., a plurality of comodulation strength values mk, or a plurality of comodulation coefficients ck; e.g., respective comodulation strength values mkassociated with different respective frequency ranges; e.g., respective comodulation coefficients c associated with different respective frequency ranges) describing a comodulation between different frequency ranges of the input audio signal (e.g., describing a comodulation of temporal envelopes in neighboring or distant (space) frequency bands) (e.g., describing a comodulation of an envelope signal associated with a currently considered frequency range, e.g., having a frequency range index k, with one or more other envelope signals associated with one or more different frequency ranges, e.g., having frequency range indices u and / or I) (e.g., on the basis of the envelope information).

[0217] The method further comprises applying (e.g., in a per frequency-range manner) a post¬ masking modeling (e.g., “pm") to the envelope information (e.g., efc(n)), or to a pre-processed version thereof (e.g., to an averaged and down-sampled version thereof; e.g., to ek(m)), to obtain a post-masking-processed envelope information (e.g., pfe(m)), which describes a masking threshold (and which may, for example, serve as a masking threshold information).

[0218] The method also comprises adapting a post-masking decay time of the post-masking modeling (e.g., a post masking decay time constant or decay factor dkused in the post masking modeling) (e.g., in a per-frequency-range manner) in dependence on the comodulation strength information (e.g., in dependence on a comodulation strength value mk(n) or in dependence on a comodulation coefficient ck) (e.g., in dependence on a comodulation strength value mk( ) associated with a currently considered frequency range, or in dependence on a comodulation coefficient ckassociated with a currently considered frequency range) (e.g., in order to obtain a faster reduction of a post-masking effect in the presence of a comparatively stronger comodulation when compared to the presence of a comparatively weaker comodulation).

[0219] This embodiment is based on the finding that an adaptation of the post masking decay time of the post masking modeling in dependence on the comodulation strength information is well adapted to characteristics of the human hearing and therefore results in masking threshold values of good quality without requiring excessive computational complexity. For example, the concept allows to have a relatively short decay time of the post masking modeling in the case of a comparatively high comodulation between different frequency ranges, and to have a comparatively long decay time of the post masking modeling in case

[0220] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 27

[0221] of a comparatively low comodulation between different frequency ranges, which has been found to well reflect the behavior of the human auditory system. This is due to the recognition that this concept is well suited to reflect or model the so-called dip listening effect of the human auditory system.

[0222] Worded yet differently, in this embodiment, the decay time of the post masking modeling may be adjusted in such a manner that it is well adapted to the human hearing, since the concept takes into account a variability of the decay time of the post masking modeling and the dependency of the decay time of the post masking modeling from the comodulation. As a consequence, the resulting masking threshold information, which is obtained using the application of the post masking modeling with adjustable decay time is particularly well adapted to an actual masking threshold occurring in a human auditory system.

[0223] Furthermore, it has been recognized that both the derivation of the decay time of the post masking modeling and the realization of a post masking modeling using the variable decay time can be implemented with comparatively low complexity, e.g. since the decay time may be derived from a value (e.g. from a single scalar value) representing a comodulation strength using a scalar (and possibly memory-less) mapping function of relatively low complexity, e.g. using a relatively simple linear or piecewise-linear mapping function. Furthermore, it has also been recognized that the application of the post masking modeling with variable decay time can also be implemented with low computational complexity, e.g. using a simple recursive mapping approach which only becomes effective for a reduction (or significant reduction) of an (e.g. preliminary) masking threshold value which is input into the post-masking modeling.

[0224] To conclude, the embodiment described is well suited to provide masking threshold values of high quality while keeping a computational complexity reasonably small.

[0225] An embodiment according to the present invention creates a method for providing a quantization step size information (qi, qj) on the basis of an input audio signal, wherein the method comprises providing a masking threshold information, according to the previously mentioned method of a providing a masking threshold information, wherein the method comprises determining a masking threshold information (e.g., pkm)) on the basis of an input audio signal (e.g., x(n)),

[0226] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX ...

[0227] 28

[0228] wherein the method comprises determining the quantization step size information using the masking threshold information (e.g., using a post-masking-processed envelope information).

[0229] This embodiment of the invention is based on the findings, that utilizing a masking threshold 5 determination method according to embodiments of this invention in order to obtain a quantization step size information can lead to a high quality quantization step size determination with high computational efficiency.

[0230] An embodiment according to the invention creates a method for providing an encoded audio 10 representation on the basis of one or more input audio signals,

[0231] wherein the method comprises providing a quantization step size information according to the previously mentioned method of providing a quantization step size information, wherein the providing the quantization step size information comprises obtaining the quantization step size information on the basis of the one or more input audio signals. 15 The method further comprises quantizing different frequency ranges of the one or more input audio signals (e.g., spectral domain coefficients (e.g., MDCT coefficients) representing the different frequency ranges, or sub-band signals representing the different frequency ranges) using different effective quantization resolutions (e.g., using a frequency-dependent scaling, to obtain a scaled representation of the different frequency ranges of the one or 20 more input audio signals, and using a quantization of the scaled representation of the different frequency ranges, wherein the frequency-dependent scaling is controlled in dependence on the quantization step size information provided by the quantization step size determinator), to obtain quantized representations of the different frequency ranges of the one or more input audio signals, and

[0232] 25 the method comprises encoding (e.g., to losslessly encode) the quantized representations of the different frequency ranges of the one or more input audio signals, to obtain the encoded representation of the one or more input audio signals.

[0233] This embodiment of the invention is based on the findings, that utilizing a quantization step size determination method according to embodiments of this invention in order to obtain an 30 encoded audio representation can lead to a higher perceptual audio quality of the one or more encoded audio signal with the same advantages as mentioned earlier.

[0234] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX An embodiment according to an invention creates a computer program for performing any of the two previously explained methods when the computer program runs on a computer. This embodiment of the invention is based on the findings, that applying the previously explained methods on a computer can be an efficient way of reaching an improved compromise between a quality of a determination of a masking threshold, an implementation complexity and possibly an achievable perceptual audio quality.

[0235] BRIEF DESCRIPTION OF THE FIGURES

[0236] Embodiments of the present invention are described herein in detail with respect to the appended drawings and figures, in which:

[0237] Fig. 1 shows a schematic illustration of a masking threshold determinator according to embodiments of the inventive concept;

[0238] Fig. 2 shows a detailed schematic illustration of a masking threshold determinator according to embodiments of the inventive concept;

[0239] Fig. 3 shows a schematic illustration of a quantization step size determinator, encompassing a masking threshold determinator, according to embodiments of the inventive concept;

[0240] Fig. 4 shows a schematic illustration of a quantization step size determinator, encompassing a masking threshold determinator with a comodulation masking release, according to embodiments of the inventive concept; Fig. 5 shows a schematic illustration of a quantization step size determinator, encompassing a masking threshold determinator with comodulation controlled post-masking, according to embodiments of the inventive concept;

[0241] Fig. 6 shows a schematic illustration of an experiment setup for investigating comodulation masking release;

[0242] Fig. 7 shows the results of the experiment depicted in Fig. 6;

[0243] Fig. 8 shows the results of a listening test comparing the comodulation controlled post-masking according to embodiments of the inventive

[0244] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 30

[0245] concept to another approach without comodulation controlled postmasking;

[0246] Fig. 9 shows a flowchart of a method for providing a masking threshold information;

[0247] Fig. 10 shows a schematic illustration of an encoder according to embodiments of the inventive concept.

[0248] DETAILED DESCRIPTION OF THE FIGURES

[0249] In the following description, embodiments are discussed in detail, however, it should be appreciated that the embodiments provide many applicable concepts that can be embodied in a wide variety. The specific embodiments discussed are merely illustrative of specific ways to implement and use the present concept, and do not limit the scope of the embodiments. In the following description of embodiments, the same or similar elements or elements that have the same functionality are provided with the same reference sign or are identified with the same name, and a repeated description of elements provided with the same reference number or being identified with the same name is typically omitted. In the following description, a plurality of details is set forth to provide a more thorough explanation of embodiments of the disclosure.

[0250] However, it will be apparent to one skilled in the art that other embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown in diagram form rather than in detail in order to avoid obscuring examples described herein. In addition, features of the different embodiments described herein may be combined with each other, unless specifically noted otherwise.

[0251] Embodiments according to Fig. 1

[0252] In the following, a masking threshold determinator 100 will be described taking reference to Fig- 1-

[0253] The masking threshold determinator 100 according to Fig. 1 is configured to receive an input audio signal 110 and to provide, on the basis thereof, a masking threshold information 112. The masking threshold determinator 100 comprises an envelope information

[0254] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 31

[0255] determination / envelope information determinator 120 which is configured to obtain envelope information 122 describing envelopes of different frequency ranges of the input audio signal 110. The masking threshold determinator 100 further comprises a comodulation strength information determination / comodulation strength information 5 determinator 130 which is configured to determine a comodulation strength information 132 describing a comodulation between different frequency ranges of the input audio signal. For example, the comodulation strength information determination / comodulation strength information determinator 130 may receive the input audio signal 110 or the envelope information 122. However, the comodulation strength information 10 determination / comodulation strength information determinator 130 may alternatively use different input information to derive the comodulation strength information 132.

[0256] Moreover, the masking threshold determinator 100 may comprise a post-masking modeling application / post-masking modeling applicator 140 which is configured to apply a post¬ 15 masking modeling to the envelope information 122, or to a preprocessed version / preprocessed envelope information 124 of the envelope information 122, to obtain a post-masking processed envelope information 142. For example, the preprocessed envelope information 124 may be obtained using a (optional) preprocessing / preprocessor 125 which is configured to apply a preprocessing to the envelope information 122. For 20 example, the post-masking process envelope information 142 may represent the masking threshold information 112, or the masking threshold information 112 may be obtained on the basis of the post-masking processed envelope information 142.

[0257] Moreover, the masking threshold determinator 100 may comprise a post-masking decay 25 time adaptation 150, wherein the post-masking decay time adaption 150 is configured to adapt a post-masking decay time of the post-masking modeling application / post-masking modeling applicator 140 in dependence on the comodulation strength information 132. For example, the post-masking decay time adaptation 150 may receive the comodulation strength information 132 and provide on the basis thereof a control information or decay 30 time information 152, which defines an adaptation of the post-masking decay in the postmasking modeling application / post-masking modeling applicator 140.

[0258] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 32

[0259] However, it should be noted that the functionalities of the masking threshold determinator 100, which have been described above, may be implemented using different functional blocks, such that the functional blocks mentioned above should be considered as examples only for the implementation of the functionality described here.

[0260] Regarding the functionality of the masking threshold determinator 100, it should be noted that the masking threshold determinator 100 derives the masking threshold information using a post-masking modeling, which is applied to the envelope information 122 or to a preprocessed version 124 of the envelope information 122, and that a post-masking decay time that is used by the post-masking modeling is adjusted in dependence on the comodulation strength (represented, for example, by the comodulation strength information 132). Accordingly, the determination of the masking threshold information 112 can be adapted to the characteristics of the human auditory system, since it has been recognized that a comodulation between different frequency ranges of an input audio signal affects a post-masking characteristic of the human auditory system. Consequently, in the masking threshold determinator 100, such a variation of the post-masking characteristic (e.g., of the post-masking decay time) may be modeled by exploiting the determination of the comodulation strength information 132. Consequently, the masking threshold information 112 is obtained with a good quality while keeping the computational complexity reasonably small.

[0261] Moreover, it should be noted that the masking threshold determinator 100 may optionally be supplemented by any of the features, functionalities and details disclosed herein. Moreover, any of the features, functionalities and details disclosed with respect to the masking threshold determinator 100 may optionally be introduced into any of the other embodiments disclosed herein, both individually and taken in combination.

[0262] Embodiment according to Fig. 2

[0263] In the following, a masking threshold determinator 200 will be described taking reference to Fig. 2.

[0264] The masking threshold determinator 200 according to Fig. 2 is configured to receive an input audio signal 210 and to provide, on the basis thereof, a masking threshold information 280.

[0265] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX ,

[0266] The masking threshold determinator 200 comprises a filterbank 215 which is configured to obtain K frequency band value sequences (e.g. in the form of a plurality of time sequences of frequency band values), e.g., frequency band values 211 for frequency band k, wherein k is the current (e.g. currently considered) frequency band, frequency band values 213 for the frequency range I, wherein I may be defined as I = k - d, and frequency band values 212 for the frequency band u, wherein u may be defined as u = k + d2, wherein each frequency band value sequence 211, 212 and 213 is derived using a designated filter for a respective frequency range of frequency band. For example, the filter hk(n) 216 derives the frequency band values 211 (which may be considered as a frequency band information) from the audio signal 210, the filter hun) 217 that derives the frequency band values 212 (which may be considered as another frequency band information) from the audio signal 210 and the filter hi( ) 218 that derives the frequency band values 213 (which may be considered as yet another frequency band information) from the audio signal 210. It is to be noted, that according to an embodiment of this invention it can be

[0267]

[0268] = d2. For example, the filter can be real valued or complex valued, and the filter can have overlapping filter characteristics between neighboring filters.

[0269] Moreover, the masking threshold determinator 200 may comprise a plurality of (respective) magnitude determinators, which are configured to determine the respective envelope information from the respective frequency band information. For example, the magnitude determinator 221 for frequency band k, determines the envelope information 226 (e.g.; efc(n)) for the frequency band k on the basis of the frequency band information 211 of the frequency band k, the magnitude determinator 222 for the frequency band u, determines the envelope information 227 for the frequency band u (e.g.; eu(n)) from the frequency band information 212 of the frequency band u

[0270] and / or the magnitude determinator 223 for the frequency band I, determines the envelope information 228 for the frequency band I (e.g.; e?(n)) from the frequency band information 213 of the frequency band I.

[0271] Moreover, the masking threshold determinator 200 may comprise a covariance estimator 240, which is configured to determine a covariance information 244 on the basis of a plurality of envelope information items, e.g., on the basis of the envelope information 226, the envelope information 227 and / or the envelope information 228. For example, an averaging 241 (e.g., a moving average; e.g., an HR lowpass filter; e.g. a separate averaging

[0272] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX for the different envelope information items) is applied on the envelope information 226, 227 and 228 to obtain (respective) average values 243 (e.g., ^(n)

[0273]

[0274] i E exemplary). For example, the covariance information 244 is obtained using a combination / combiner 242 combining the envelope information items 226, 227 and 288 and the respective average values 243. For example, the average values 243 could be calculated as moving averages with a possible look-ahead of D samples and Length L with a

[0275]

[0276] ^n) = | Emtn+D-L+iei(m) or using first order HR lowpass filters a^n) = patn - 1) + (1 - p^e^ ) where the decay factor p with 0 < p < 1 determines the time constant and also can depend on band index i. Furthermore, the covariance information 244 may, for example, be calculated using Cfei(n) = eft(n)e;(n) - afc(n)czz(n) and Cku(n) = efc(n)eu(n) - afe(n)au(n), wherein

[0277]

[0278] the (77) denotes a short time average. For example, in case of using a first order HR lowpass filter, different decay factors for averaging envelopes or products thereof may be used.

[0279] Moreover, the masking threshold determinator 200 may comprise an optional normalization / combination determinator 245 which may be configured to obtain the comodulation strength information 246 (e.g., a normalized and optionally combined covariance information; e.g., C^, e.g., mfe(n)) using a normalization (e.g. of one or more covariance information values, or of a combined and / or preprocessed value obtained on the basis of one or more covariance values) and optionally also using a combination (e.g. of a plurality of covariance information values, or of a plurality of preprocessed covariance information values).

[0280] Moreover, the masking threshold determinator 200 may comprise a mapping / mapper 250 which is configured to apply one or more mapping operations, e.g., a sequence of mapping operations, (or, equivalently, a combined mapping operation) to the comodulation strength information 246 (e.g. to a value of the comodulation strength information), to obtain the decay factor 256, (e.g., dk). For example, the comodulation strength information 246 (e.g., a value of the comodulation strength information) is mapped onto a comodulation coefficient 252 (e.g., cfc(n); e.g., (n)), using a mapping 251 onto a comodulation coefficient 252. For example, the comodulation coefficient 252 is mapped onto a decay time constant value 254 (e.g., varDecayTimeConst) using a mapping to decay time constant 253. For example, the decay factor 256 is obtained on the basis of the decay time constant value 254 using a mapping to decay factor for post masking modeling application 255.

[0281] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0282] C / For example, the mappings 251 and 253 may be substantially linear mappings, e.g., with a saturation limiting a range of values. For example, the mapping 255 may comprise a combination of a constant term (e.g. 1) with a negative hyperbolic mapping of the decay time constant value 254.

[0283] Moreover, the masking threshold determinator 200 may comprise an optional averaging 230 which is configured to average the envelope information k 226 for frequency band k to obtain the averaged envelope information 229 for frequency band k. Furthermore, the averaged envelope information 229 (or, alternatively, the envelope information k 226 for frequqncy band k) may be further processed using an optional downsampling 231 to obtain a pre-processed (e.g. down-sampled) envelope information 224 (e.g., ek' m)).

[0284] Moreover, the masking threshold determinator 200 may comprise an optional attenuation factor determinator 260 which is configured to apply an attenuation factor determination to the covariance information 244, or to the comodulation strength information 246, or to the comoduiation coefficient 252, in order to obtain an attenuation factor 261 (e.g., gk(m)-, e.g., gk(m)). In other words, generally speaking, the optional attenuation factor determination may determine the attenuation factor 261 in dependence on (or on the basis of) the covariance information 244. For example, the attenuation factor is used to optionally scale the pre-processed envelope information e (n) using an optional scaling / scaler 262 to obtain a scaled pre-processed envelope information 264.

[0285] For example, the dashed method blocks (or method steps) are a variant (or are part of a variant) in which the CMR estimation disclosed herein (or, optionally a different CMR estimation) is used in already published methods (see, for example, [3] and / or [4]) for a direct reduction of the masking thresholds. Worded differently, the dashed method blocks or method steps, e.g. method blocks or method steps 260,262,263, may, for example, perform a scaling as described in [3] and / or in [4], wherein, for example, one or more of the CMR estimation concepts disclosed herein may be used, e.g. to determine the comoduiation information used to adapt the scaling.

[0286] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0287] C / ! Moreover, the masking threshold determinator 200 may comprise a post masking modeling application 270 which is configured to obtain a post-masking processed envelope information 271 on the basis of the pre-processed envelope information 224 (e.g., ek(m)) or on the basis of the scaled pre-processed envelope information 264 (or on the basis of 5 the envelope information 226, or on the basis of the averaged envelope information 229) and in dependence on the decay factor 256.

[0288] For example, the post-masking modeling application 270 may apply an envelope following with an adjustable decay time to a sequence of input values of the post-masking modeling 10 application, wherein the adjustable decay time is determined by the decay factor 256. As an example, the post-masking modeling application may apply a max operation 272 to obtain the post masking processed envelope information 271 (e.g., Pfc(m)) on the basis of a decay factor 256 and a scaled pre-processed envelope information 264 (or, alternatively, the pre-processed envelope information 224, or, alternatively, the averaged envelope 15 information 229, or, alternatively, the envelope information 226).

[0289] For example, the post-masking modeling application 270 is configured to obtain an adjusted “previous” post masking processed envelope information 274 (to be precise, a value of the adjusted previous post masking processed envelope information 274) by applying a product 20 operation (or multiplication operation, or scaling operation) 273 to the decay factor 256 and the post masking processed envelope information 271 of a previous timestep (e.g., pk(m - 1)) (to be precise, to a previous value of the post masking processed envelope information 271), for example a product operation dkpkm - 1).

[0290] For example, the max operation 272 is configured to obtain the post masking processed 25 envelope information 271 (e.g. a current value pk(m) of the post masking processed envelope information 271) using a maximum operator, which is applied to the pre-processed envelope information 224 (e.g. to a current value e (m) of the scaled pre-processed envelope information 264) or to the scaled pre-processed envelope information 264 and to the adjusted previous post masking processed envelope information 274 (e.g., to a previous 30 value dkPk(m-1) of the adjusted previous post masking processed envelope information 274), and provides, as a result, the current value (e.g. pk(m)) of the post masking processed envelope information 271.

[0291] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0292] C / 1 Moreover, the masking threshold determinator 200 may comprise an optional scaling / scaler 263 to apply a scaling operation to the post masking processed envelope information 271 to obtain the masking threshold information 280.

[0293] Alternatively, however, the post masking processed envelope information 271 may serve as the masking threshold information 280.

[0294] To conclude, the masking threshold determinator 200 may extract the envelope information 226 on the basis of the input audio signal 210. In a first topologically linear processing path, comprising, for example, the (optional) averaging 230, the (optional) downsampling 231, the (optional) scaling 262, the post-masking modeling application 270 and the (optional) scaling 263, an effect of a post-masking may be applied to a (optionally averaged and / or down-sampled) sequence of envelope values associated with a currently considered frequency band (e.g. having band index k). A decay time of the post masking modeling, which is applied to the (optionally averaged and / or down-sampled) sequence of envelope values, is determined in a second processing path, which comprises a determination of the comodulation strength information 246 (e.g. of a comodulation strength information describing the comodulation between the currently considered frequency band and one or more other frequency bands) (e.g. using functional blocks 240 and 245) and a derivation of the decay factor 256 from the comodulation strength information 246. The decay factor 256 effectively determines a decay time of the post-masking modeling application 270, and therefor determines how fast a value of the post-masking processed envelope information 271 decays in response to a comparatively fast (e.g. step-wise) reduction of the sequence of (optionally averaged and / or down-sampled) envelope information values.

[0295] Thus, as a consequence, an intensity of a comodulation of the currently considered frequency band (e.g., having frequency band index k) with one or more other frequency bands (e.g., having frequency band indices I and u) determines a decay characteristic of the post masking modeling application.

[0296] The optional scaling 262 or 263 may further effect an adjustment of the magnitude of the masking threshold information in dependence on the intensity of a comodulation of the currently considered frequency band (e.g. having frequency band index k) with one or more

[0297] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0298] C / 1 38

[0299] other frequency bands (e.g. having frequency band indices I and u) and even improve the quality of the masking threshold information.

[0300] It is to be noted, that even though in the Fig. 2 three frequency bands are depicted, this merely represents an exemplary illustration, wherein according to embodiments of the invention the masking threshold determinator 200 can utilize one or more frequency bands with no restriction on the number of frequency bands.

[0301] Moreover, it should be noted that the masking threshold determinator 200 may optionally be supplemented by any of the features, functionalities and details disclosed herein. Moreover, any of the features, functionalities and details disclosed with respect to the masking threshold determinator 200 may optionally be introduced into any of the other embodiments disclosed herein, both individually and taken in combination.

[0302] Embodiment according to Fig. 3

[0303] In the following a quantization step size determinator 300 will be described taking reference to Fig. 3. The Fig. 3 may, for example, also be describes as “Simple filter bank based perceptual model"

[0304] It is to be noted, that all embodiments disclosed herein, even though described in the context of a quantization step size determinator, can all be associated with a masking threshold determinator (wherein, for example, the block 370 is not required) or an audio encoder (wherein, fore example, an encoding functionality using the quantization step size information is added), according to embodiments of this invention.

[0305] The quantization step size determinator according to Fig. 3 is configured to receive an input audio signal 310 and provide, on the basis thereof, a quantization step size information, e.g. comprising a quantization step size value 311 and a quantization step size value qj 312. For example, the quantization step size determinator may comprise a masking threshold determinator 301 and a mapping and threshold in quiet determination / determinator 370. For example, the masking threshold determinator 301 is configured to provide, on the basis of the input audio signal 301, one or more masking thresholds (e.g masking threshold values

[0306] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX or sequences of masking threshold values), exemplary depicted here with the masking threshold 366 for the frequency band u, the masking threshold 367 for the frequency band k and the masking threshold 368 for the frequency band I. For example, the mapping and threshold in quiet determination / determinator 370 is configured to obtain the quantization 5 parameter (e.g quantization step size parameter) qt 311 and the quantization parameter (e.g quantization step size parameter) qj 312 on the basis of one or more masking thresholds (e.g masking threshold values or sequences of masking threshold values), exemplary depicted here with the masking threshold 366 of for frequency band u, the masking threshold 367 for the frequency band k and the masking threshold 368 for the 10 frequency band I,

[0307] Moreover, the masking threshold determinator 301 may be configured to process the exemplary frequency bands u, k and I, wherein for each frequency band the masking threshold determinator 301 is configured to obtain a masking threshold, exemplary depicted 15 here with the masking threshold 366 for the frequency band u, the masking threshold 367 k for the frequency band and the masking threshold 368 for the frequency band I. For example, for each frequency band the masking threshold may be obtained by applying a frequency filter (e.g. of a filterbank) to the input audio signal 310, exemplary depicted here with the filter hu(n) 321 for the frequency bandu, the filter hk(n) 322 for the frequency band 20 k and the filter (n) 323 for the frequency band I, to obtain the frequency band information items (e.g. in the form of sequences of real-valued or complex-valued frequency band values), exemplary depicted here with the frequency band information 326 for the frequency band u, the frequency band information 327 for the frequency band k and the frequency band information 328 for the frequency band I.

[0308] 25

[0309] Moreover, the masking threshold determinator 301 may comprise a plurality of magnitude determinators (|.|), exemplary depicted here with the magnitude determinator 331 for the frequency band u, the magnitude determinator 332 for the frequency band k and the magnitude determinator 333 for the frequency band I, that are configured to determine 30 envelope information items, exemplary depicted here with the envelope information eM(n) 336 for the frequency band u, the envelope information ek(n) 337 for the frequency band k and the envelope information et(n) 338 for the frequency band I, respectively, on the basis of the frequency band information items, exemplary depicted here with the frequency

[0310] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0311] Gn information 326 for the frequency band u, the frequency information 327 for the frequency band k and the frequency information 328 I, respectively.

[0312] Moreover the masking threshold determinator 301 may comprise a plurality of average 5 determinators (“avg”), that are configured to obtain average envelope information items by averaging the envelope information items, exemplary depicted her with the average determinator 341 for the frequency band u which is configured to obtain the average envelope information 346 for the frequency band u from the envelope information 336 of the frequency band u, the average determinator 342 for the frequency band k which is 10 configured to obtain the average envelope information 347 for the frequency band fcfrom the envelope information 337 of the frequency band k, and the average determinator 343 for the frequency band I which is configured to obtain the average envelope information 348 for the frequency band I from the envelope information 338 of the frequency band I.

[0313] 15 Moreover, the masking threshold determinator 301 may comprise a plurality of down samplers (shown as I P) that are configured to downsample the respective average envelope information to obtain the respective preprocessed envelope information; the down-samplers are exemplary depicted here with the downsampler 351 for the frequency band u, which is configured to obtain the preprocessed envelope information 356 (e.g., 20 e„(n)) for the frequency band u on the basis of the average envelope information 346 of the frequency band u, with the downsampler 352 for the frequency band k which is: configured to obtain the preprocessed envelope information 357 (e.g., e^(n)) for the frequency band k on the basis of the average envelope information 347 of the frequency band k, and with the downsampler 353 for the frequency band I which is configured to 25 obtain the preprocessed envelope information 358 (e.g., e((n)) for the frequency band I on the basis of the average envelope information 348 of the frequency band I.

[0314] i Moreover, the masking threshold determinator 301 may comprise a plurality of post¬ masking modelling applications (shown as pm), that are configured to determine the: 30 respective masking thresholds on the basis of the respective preprocessed envelope: information; the post-masking-modeling applications are exemplary depicted here with the post-masking modelling application 361 for the frequency band u which is configured to determine the masking threshold 366 for the frequency band u on the basis of the

[0315] i FHIIS24EM24 FHIIS24EM24-2024346428. DOCX preprocessed envelope information 356 of the frequency band u 356, the post-masking modelling application 362 for the frequency band k which is configured to determine the masking threshold 367 for the frequency band k on the basis of the preprocessed envelope information 357 of the frequency band k and the post-masking modelling application 363 for the frequency band I which is configured to determine the masking threshold 368 for the frequency band I on the basis of the preprocessed envelope information 358 of the frequency band I.

[0316] It is to be noted, that even though in the Fig. 3 three frequency bands are depicted, this merely represents an exemplary illustration, wherein according to embodiments of the invention the masking threshold determinator 301 can utilize one or more frequency bands with no restriction on the number of frequency bands.

[0317] Moreover, it should be noted that the masking threshold determinator 301 may optionally be supplemented by any of the features, functionalities and details disclosed herein. Moreover, any of the features, functionalities and details disclosed with respect to the masking threshold determinator 301 may optionally be introduced into any of the other embodiments disclosed herein, both individually and taken in combination.

[0318] In particular, it should be noted that, in some embodiments, decay times of the post masking modeling applications 361, 362, 363 may optionally be adapted in dependence on comodulation characteristics of respective envelope signals, as disclosed herein.

[0319] Embodiments according to Fig. 4

[0320] In the following a quantization step size determinator 400 will be described taking reference to Fig. 4. The Fig.4 may, for example, be also describes as “Perceptual model with direct reduction of masking thresholds controlled by comodulation strength”.

[0321] For example, the quantization step size determinator 400 (or, generally speaking the concept or method as described with respect to Fig. 4) is a variant (or are part of a variant) in which the CMR estimation disclosed herein (or, optionally a different CMR estimation) is used in an already published method, or in in already published methods (see, for example, [3] and / or [4]), for a direct reduction of the masking thresholds.

[0322] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX Worded differently, the quantization step size determinator 400 may, for example, perform a scaling which is similar or even equal to the concept as described in [3] and / or in [4], wherein, for example, one or more of the CMR estimation concepts disclosed herein may be used, e.g. to determine the comodulation information used to adapt the scaling.

[0323] It is to be noted, that ail embodiments disclosed herein, even though described in the context of a quantization step size determinator, can all be associated with a masking threshold determinator (wherein, for example, the block 370 is not required) or an audio encoder (wherein, fore example, an encoding functionality using the quantization step size information is added), according to embodiments of this invention.

[0324] The quantization step size determinator according to Fig. 4 is configured to receive an input audio signal 410 and provide, on the basis thereof, a quantization step size information, e.g. comprising a quantization step size value qt411 and a quantization step size value qj 412. For example, the quantization step size determinator may comprise a masking threshold determinator 401 and a mapping and threshold in quiet determination / determinator470. For example, the masking threshold determinator 401 is configured to provide, on the basis of the input audio signal 401, one or more masking thresholds (e.g. masking threshold values or sequences of masking threshold values), exemplary depicted here with the masking threshold 466 for the frequency band u, the masking threshold 467 for the frequency band k and the masking threshold 468 for the frequency band I. For example, the mapping and threshold in quiet determination / determinator 470 is configured to obtain the quantization parameter (e.g. quantization step size parameter) qt411 and the quantization parameter (e.g. quantization step size parameter) qj 412 on the basis of one or more masking thresholds (e.g. masking threshold values or sequences of masking threshold values), exemplary depicted here with the masking threshold 466 of for frequency band u, the masking threshold 467 for the frequency band k and the masking threshold 468 for the frequency band I.

[0325] Moreover, the masking threshold determinator 401 may be configured to process the exemplary frequency bands u, k and I, wherein for each frequency band the masking threshold determinator 401 is configured to obtain a masking threshold, exemplary depicted here with the masking threshold 466 for the frequency band n, the masking threshold 467

[0326] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX k for the frequency band and the masking threshold 468 for the frequency band I. For example, for each frequency band the masking threshold may be obtained by applying a frequency filter (e.g. of a filterbank) to the input audio signal 410, exemplary depicted here with the filter hu(n) 421 for the frequency ban u, the filter hk(n) 422 for the frequency band k and the filter hi(n) 423 for the frequency band I, to obtain the frequency band information items (e.g. in the form of sequences of real-valued or complex-valued frequency band values), exemplary depicted here with the frequency band information 426 for the frequency band u, the frequency band information 427 for the frequency band k and the frequency band information 428 for the frequency band I.

[0327] Moreover, the masking threshold determinator 401 may comprise a plurality of magnitude determinators (|.|), exemplary depicted here with the magnitude determinator 431 for the frequency band u, the magnitude determinator 432 for the frequency band k and the magnitude determinator 433 for the frequency band I, that are configured to determine envelope information items, exemplary depicted here with the envelope information eu(n) 436 for the frequency band u, the envelope information ek(n) 437 for the frequency band k and the envelope information ez(n) 438 for the frequency band I, respectively, on the basis of the frequency band information items, exemplary depicted here with the frequency information 426 for the frequency band u, the frequency information 427 for the frequency band k and the frequency information 428 I, respectively.

[0328] Moreover the masking threshold determinator 401 may comprise a plurality of avg determinators, that are configured to obtain respective average envelope information items by averaging the respective envelope information items, exemplary depicted here with the avg determinator 441 for the frequency band u which is configured to obtain the average envelope information 446 for the frequency band u from the envelope information 436 of the frequency band it, the avg determinator 442 for the frequency band k which is configured to obtain the average envelope information 447 for the frequency band k from the envelope information 437 of the frequency band k and the avg determinator 443 for the frequency band I which is configured to obtain the average envelope information 448 for the frequency band I from the envelope information 437 of the frequency band I.

[0329] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX Moreover, the masking threshold determinator 401 may comprise a plurality of down samplers (shown as i P) that are configured to downsample the respective average envelope information to obtain the respective preprocessed envelope information; the down-samplers are exemplary depicted here with the downsampler 451 for the frequency bandit which is configured to obtain the preprocessed envelope information 456 (e.g.,eu n)) for the frequency band u on the basis of the average envelope information 446 of the frequency band u, with the downsampler 452 for the frequency band k which is configured to obtain the preprocessed envelope information 457 (e.g., e^(n)) for the frequency band k on the basis of the average envelope information 447 of the frequency band k and with the downsampler 453 for the frequency band I which is configured to obtain the preprocessed envelope information 458 (e.g., e;'(n)) for the frequency band I on the basis of the average envelope information 448 of the frequency band I.

[0330] Moreover, the masking threshold determinator 401 may comprise a plurality of post¬ masking modelling application blocks (shown as pm), that are configured to determine the respective masking thresholds on the basis of the respective preprocessed envelope information; the post masking modeling applications (application blocks) are exemplary depicted here with the post-masking modelling application 461 for the frequency band u which is configured to determine the masking threshold 466 for the frequency band u on the basis of the preprocessed envelope information 456 of the frequency band u, the postmasking modelling application 462 for the frequency band k which is configured to determine the masking threshold 467 for the frequency band k on the basis of the preprocessed envelope information 457 of the frequency band k and the post-masking modelling application 463 for the frequency band I which is configured to determine the masking threshold 468 for the frequency band I on the basis of the preprocessed envelope information 458 of the frequency band I.

[0331] Moreover, the masking threshold determinator 401 may comprise a comodulation masking release determination / determinator480 (CMR), that is configured to obtain an attenuation factor 481 (e.g., gk(n)) on the basis of the envelope information 436 of the frequency band u, the envelope information 437 of the frequency band k and the envelope information 438 of the frequency band I. For example, the atenuation factor 481 may be utilized by a product operation or scaling operation 482 to obtain the masking threshold 467 for the frequency band k.

[0332] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 45

[0333] Accordingly, the masking threshold determinator may scale the output signals of the post masking modeling application (or, alternatively, input signals of the post masking modeling application), to thereby consider an impact of the comodulation onto the masking threshold information. For example, the product operation or scaling operation 482 may result in a scaling of a steady-state value of the masking threshold information which is determined by the comodulation (such that, for example, a comparatively high comodulation results in a comparatively small scaling factor, and consequently in a comparatively small masking threshold value). However, in contrast to the post-masking modeling, the scaling does not affect dynamic properties (e.g. a variation over time) of the masking threshold values in the case of a constant comodulation.

[0334] It is to be noted, that even though in the Fig. 4 three frequency bands are depicted, this merely represents an exemplary illustration, wherein according to embodiments of the invention the masking threshold determinator 401 can utilize one or more frequency bands with no restriction on the number of frequency bands. Furthermore, it is to be noted, that even though the comodulation masking release 480 is depicted here only for the frequency band k, embodiments of the invention by no means restrict the application of a comodulation masking release on one frequency band. Rather it was chosen to visualize the application of the comodulation masking release 480 on only one frequency band due to visualization limitations.

[0335] Moreover, it should be noted that the masking threshold determinator 401 may optionally be supplemented by any of the features, functionalities and details disclosed herein. Moreover, any of the features, functionalities and details disclosed with respect to the masking threshold determinator 401 may optionally be introduced into any of the other embodiments disclosed herein, both individually and taken in combination.

[0336] In particular, it should be noted that, in some embodiments, decay times of the post masking modeling applications 461, 462, 463 may optionally be adapted in dependence on comodulation characteristics of respective envelope signals, as disclosed herein.

[0337] Embodiment according to Fig, 5

[0338] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX In the following a quantization step size determinator 500 will be described taking reference to Fig. 5. The Fig.5 may also be describes as “Simple filter bank based perceptual model controlling postmasking decay dkby comodulation strength”

[0339] It is to be noted, that all embodiments disclosed herein, even though in the context of a quantization step size determinator can all be associated with a masking threshold determinator (wherein, for example, the block 570 is not required) or an audio encoder (wherein, for example, an encoding functionality using the quantization step size information is added), according to embodiments of this invention.

[0340] The quantization step size determinator according to Fig. 5 is configured to receive an input audio signal 510 and provide, on the basis thereof, a quantization step size information, e.g., comprising a quantization step size value qt511 and a quantization step size qj 512. For example, the quantization step size determinator may comprise a masking threshold determinator 501 and a mapping and threshold in quiet determination / determinator 570. For example, the masking threshold determinator 501 is configured to provide, on the basis of the input audio signal 501, one or more masking thresholds (e.g. masking threshold values or sequences of masking threshold values), exemplary depicted here with the masking threshold 566 for the frequency band u, the masking threshold 567 for the frequency band k and the masking threshold 568 for the frequency band I. For example, the mapping and threshold in quiet 570 is configured to obtain the quantization parameter (e.g. quantization step size parameter) qt511 and the quantization parameter (e.g. quantization step size parameter) qj 512 on the basis of one or more masking thresholds (e.g. masking threshold values or sequences of masking threshold values), exemplary depicted here with the masking threshold 566 of for frequency band u, the masking threshold 567 for the frequency band k and the masking threshold 568 for the frequency band I.

[0341] Moreover, the masking threshold determinator 501 may be configured to process the exemplary frequency bands u, k and I, wherein for each frequency band (or at least for a plurality of frequency bands) the masking threshold determinator 501 is configured to obtain a masking threshold, exemplary depicted here with the masking threshold 566 for the frequency band u, the masking threshold 567 k for the frequency band and the masking threshold 568 for the frequency band I. For example, for each frequency band the masking

[0342] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0343] C *

[0344] 47

[0345] threshold may be obtained by applying a frequency filter (e.g. of a filterbank) to the input audio signal 510, exemplary depicted here with the filter hu(n 521 for the frequency band u, the filter hk(n) 522 for the frequency band k and the filter ht(n) 523 for the frequency band I, to obtain the frequency information items (e.g. in the form of sequences of real-valued or 5 complex-valued frequency band values), exemplary depicted here with the frequency band information 526 for the frequency band u, the frequency band information 527 for the frequency band k and the frequency band information 528 for the frequency band I.

[0346] Moreover, the masking threshold determinator 501 may comprise a plurality of magnitude 10 determinators (|.|), exemplary depicted here with the magnitude determinator 531 for the frequency band u, the magnitude determinator 532 for the frequency band k and the magnitude determinator 533 for the frequency band I, that are configured to determine envelope information items, exemplary depicted here with the envelope information eM(n) 536 for the frequency band u, the envelope information ek(n) 537 for the frequency band k 15 and the envelope information ei(n) 538 for the frequency band I, respectively, on the basis of the frequency information, exemplary depicted here with the frequency information 526 for the frequency band u, the frequency information 527 for the frequency band k and the frequency information 528 I, respectively.

[0347] 20 Moreover the masking threshold determinator 501 may comprise a plurality of average determinators, that are configured to obtain average envelope information items by averaging the envelope information items, exemplary depicted her with the average determinator 541 for the frequency band u which is configured to obtain the average envelope information 546 for the frequency band u from the envelope information 536 of 25 the frequency band u, the average determinator 542 for the frequency band k which is configured to obtain the average envelope information 547 for the frequency band k from the envelope information 537 of the frequency band k and the average determinator 543 for the frequency band I which is configured to obtain the average envelope information 548 for the frequency band I from the envelope information 538 of the frequency band I.

[0348] 30

[0349] Moreover, the masking threshold determinator 501 may comprise a plurality of down samplers (shown as I P) that are configured to downsample the respective average envelope information to obtain the respective preprocessed envelop information; the down-

[0350] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX .

[0351] 48

[0352] samplers are exemplary depicted here with the downsampler 551 for the frequency band u which is configured to obtain the preprocessed envelope information 556 (e.g., e^(n)) for the frequency band it on the basis of the average envelope information 546 of the frequency band u, with the downsampler 552 for the frequency band k which is configured to obtain the preprocessed envelope information 557 (e.g., ek(n)) for the frequency band k on the basis of the average envelope information 547 of the frequency band k and with the downsampler 553 for the frequency band I which is configured to obtain the preprocessed envelope information 558 (e.g., e ( )) for the frequency band I on the basis of the average envelope information 548 of the frequency band I.

[0353] Moreover, the masking threshold determinator 501 may comprise a (or a plurality of) comodulation masking release determination / determinator 590 which is configured to determine a decay factor 591 (e.g., dk) for the frequency band k based on the comodulation strength information of the preprocessed envelope information 536 of the frequency band u, the preprocessed envelope information 537 of the frequency band k and the preprocessed envelope information 538 of the frequency band I.

[0354] Moreover, according to embodiments of this invention the masking threshold determinator 501 may comprise a (or a plurality of) post-masking modelling application (or post-masking modeling application block) 562 which is configured to determine the masking threshold 592 for the frequency band k by applying a post-masking (or a post-masking processing) to the preprocessed envelope information 557 of the frequency band k and adapting postmasking processing in dependence on the decay factor 591.

[0355] Moreover, the masking threshold determinator 501 may comprise a plurality of postmasking modelling application blocks (shown as pm), that are configured to determine the masking thresholds (e.g. respective masking threshold values associated with respective frequency bands) on the basis of the preprocessed envelope information; the post-masking modeling applications (application blocks) are exemplary depicted here with the post-masking modelling application 561 for the frequency band u which is configured to determine the masking threshold 566 for the frequency band u on the basis of the preprocessed envelope information 556 of the frequency band u and the post-masking modelling application 563 for the frequency band I which is configured to determine the

[0356] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 49

[0357] masking threshold 568 for the frequency band I on the basis of the preprocessed envelope information 558 of the frequency band I.

[0358] To conclude, the masking threshold determinator may generate the respective masking threshold information, here depicted with the respective masking thresholds (or masking threshold values), on the basis of the respective preprocessed envelope information and utilizes the comodulation of other, e.g., neighboring frequency bands (e.g. spaced by d-1 frequency bands in between) with the currently considered frequency band in order to further adapt the provision (e.g. derivation) of the currently considered masking thresholds (e.g. in order to adapt a decay time of the post-masking modelling application 562). This is visualized for the frequency band k, wherein the preprocessed envelope information 557 of the frequency band k is obtained from (e.g. on the basis of) the frequency information 427 of the frequency band k through the usage of the magnitude determinator 432, the average determinator 442 and the downsampler 452, and wherein the comodulation information is used to determine the decay factor 591, and wherein the comodulation information is obtained based on the envelope information 537 of the frequency band k, the exemplary envelope information 536 of the frequency band u and the exemplary envelope information 538 of the frequency band I using the comodulation masking release determination / determinator 590.

[0359] In the Fig. 5. the comodulation consideration is only visualized for the frequency band k this however was merely done due to visualization limitations and by no means represents a restriction on one frequency band.

[0360] It is to be noted, that even though in the Fig. 5 three frequency bands are depicted, this merely represents an exemplary illustration, wherein according to embodiments of the invention the masking threshold determinator 501 can utilize one or more frequency bands with no restriction on the number of frequency bands. Furthermore, it is to be noted, that even though the comodulation masking release 590 is depicted here only for the frequency band k, embodiments of the invention by no means restrict the application of a comodulation masking release on one frequency band. Rather it was chosen to visualize the application of the comodulation masking release 580 on only one frequency band due to visualization limitations.

[0361] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX Moreover, it should be noted that the masking threshold determinator 501 may optionally be supplemented by any of the features, functionalities and details disclosed herein. Moreover, any of the features, functionalities and details disclosed with respect to the masking threshold determinator 501 may optionally be introduced into any of the other embodiments disclosed herein, both individually and taken in combination.

[0362] Explanation of some background based on Figs. 6, 7 and 8

[0363] In the following an experiment setup will be described taking reference to Fig. 6. The experiment is designed to show the phenomenon called comodulation masking release (CMR). Two types of maskers were used in this experiment. The first was stationary band limited noise 650 centered around 1 kHz with the bandwidth as a control parameter. And the second was noise 620 with the same bandwidth, but multiplied by a randomly, relatively slowly changing amplitude modulation function 630. The test signal 640 was a 1 kHz tone adjusted in its level, so that it just started becoming audible, as in usual masking experiments. The results of a listening test of the resulting audio signal 610 of this experiment are depicted in Fig. 7.

[0364] In the following the results in of the experiment as described in Fig.6 will be described taking reference to Fig. 7. Fig. 7 depicts a plot 710, showing the masker bandwidth in kHz on the x-axis ranging from nearly 0 kHz to 1.0 kHz and the signal threshold in dB SPL on the y- axis. Moreover, the plot encompasses two lines, the audibility thresholds of a 1 kHz test tone in dependency of noise masker bandwidth without amplitude modulation 720 and the audibility thresholds of a 1 kHz test tone in dependency of noise masker bandwidth with low frequency random modulation 730, both according to the experiment design of Fig. 6. The result shown in Fig. 7 for the audibility thresholds of a 1 kHz test tone in dependency of noise masker bandwidth without amplitude modulation 720 looks as expected: the masking threshold increases with bandwidth until the critical bandwidth is reached and then remains constant. The masking threshold starts at around 52,5 dB SPL for a masker bandwidth near zero, rises nearly linear up to just under 60 at around.15 kHz Masker bandwidth and then stays constant up until a Masker bandwidth of 1.0 kHz. The result for the audibility thresholds of a 1 kHz test tone in dependency of noise masker bandwidth with low frequency random modulation 730 however looks surprising. At first, the masking threshold increases with bandwidth, but then decreases again. The masking threshold starts at around 52 db SPL

[0365] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX and rises to around 56 dB SPL at around.1 kHz and then falls of till the end of the plot at 1 kHz to around 50 dB SPL. This means, that although the total masker level increases, the masking threshold decreases. Or, in other words, the noise components added by the bandwidth increase rather help for the audibility of the test tone than prevent it.

[0366] Hence, according to an aspect of the invention, a masking threshold may be reduced with increasing comodulation between two or more frequency ranges (or frequency bands).

[0367] For example, the scaling 262, or the scaling 263, or the scaling 482, which may optionally be introduced into any of the embodiments of the present invention, may have the effect that the masking threshold is (e.g. statically or qausi-statically) reduced with increasing comodulation between two or more frequency ranges (or frequency bands). Accordingly, a reliability of masking threshold values may be improved.

[0368] in the following the results of a listening experiment will be described taking reference to Fig. 8. Fig. 8 depicts a plot 810, showing the result of a listening test, visualized with average and 95% confidence intervals for Student’s t Distribution for eight subjects. For every subject the plot 810 depicts a result for four models, a conventional approach within an MPEG-H based transform coding framework 820, a model with a comodulation masking release according to embodiments of this invention within an MPEG-H based transform coding framework, a model based on anchor_7k and a hiddenRef model. Note, that the latter two are merely there for reference. Both first mentioned models were tuned by scaling their resulting masking thresholds in a way that they achieved equal average bit rates for a very long item concatenated from a multitude of test signals. Here this was set to 56 kbit / s. Nevertheless, the bit rates for individual items used in the test can differ. The results illustrated in Fig. 8 indicate an average improvement by applying comodulation controlled post-masking according to an embodiment of this invention.

[0369] Further aspects and embodiments

[0370] FHHS24EM24 FHIIS24EM24-2024346428. DOCX

[0371] C In the following further aspects and embodiments according to the present invention will be described, which may be used independently, and which may optionally be introduced into any of the embodiments disclosed herein, both individually and taken in combination.

[0372] The following may address in particular features, functionalities and details regarding embodiments for a Perceptual Model Considering Comodulation Masking Release by Postmasking (e.g., post-masking) Adaptation

[0373] 10 Estimation of Comodulation in Filter Bank based Perceptual Models

[0374] Assuming that a perceptual model uses K filters k (e.g., the frequency filters 216, 217, 218; 321, 322, 323; 421, 422, 423; 521, 522, 523) with 0 < k < K for approximating the spectral decomposition in the human auditory system, covariances of temporal envelopes in neighboring bands can be evaluated. For real valued filters, their outputs should 15 preferably (but not necessarily) be rectified and smoothened (e.g. by lowpass filtering).

[0375] For complex valued filters, the magnitudes of the outputs can, for example, be directly used. Differences of signal delays in the different bands may, for example, be taken into account by proper compensation. For example, in [3], cross-correlations are estimated by inner products of normalized envelopes in the corresponding bands. From those, an 20 “effective bandwidth” for each band can be obtained, for example, by averaging over cross-correlations with multiple (e.g., 5) lower and upper neighbors. The amount of masking threshold reduction can then, for example, be derived from this effective bandwidth and the ratio between the minimum and the maximum of the envelope in this band within a certain time interval (e.g., 20... 30ms). This ratio is, for example,

[0376] 25 introduced to take into account the dip listening effect. In some implementations, the overall complexity of these operations, however, is relatively high and thus they are, for example, applied with a relative low temporal resolution.

[0377] Since, for example, magnitude responses of directly neighboring bands usually (but not 30 necessarily) significantly overlap, their output magnitudes may, in some cases, show strong covariances even with strongly band limited modulated maskers. Therefore, for example, an appropriate distance d between band indices can be selected. For example, the covariance of bands k with k > d and I = k - d at sample position n can then be

[0378] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX estimated using short time averages (denoted as (...)) from the envelopes efc(n) and

[0379]

[0380] Cfcj(n) = ek(n)ei(n) - ak(n)al(n') with ak(n) = e^(nj and al(n') = e[(n).

[0381] The envelopes ek(n) and ez(n) may, for example, be envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528) and Cklmay, for example, be an example for a component of the covariance information 244.

[0382] The comodulation strength should (preferably, but not necessarily) be independent of the total strength of the signal, what could be achieved by using the correlation coefficient:

[0383] - J4(n) ef(n)

[0384]

[0385] This calculation of Cklmay be an example for the calculation of the comodulation strength information 132, 246.

[0386] Another normalization, however, was found to be even more useful, since for bands k with k > d and k < K - d, a combination of correlations with a lower neighbor I = k - d and an upper neighbor u = k + d can be used to give a better indication of comodulations. In this case, the following

[0387] combination of Ckiand Ckuand the corresponding averages can, for example, be used, since it was found to give good results:

[0388] G^ / C^nj +

[0389] m

[0390]

[0391] k n2ak(n) + at(n) + au(n)

[0392] with, e.g., G = 2.8;

[0393] or, alternatively, for example, the following combination of Ckland Ckwcan be used

[0394]

[0395] kmax (ak(n), at(n), au(n))

[0396] with, e.g., G = 0.7

[0397] where G is a normalization factor to obtain averages similar to the cases, in which only one neighbor with the desired distance d can be used. In other words, for example, instead of using 2ak+ at+ auin the denominator, the maximum of ak, atand aumay as well be used. In that case, for example, a G of 0.7 may be used.

[0398] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX In those cases, where either no lower or no upper neighbor is available, another normalized covariance with the upper or lower neighbor can be used, i.e.:

[0399] if k < d

[0400] mk(n) with j =

[0401]

[0402] ak(n) + ay(n) if k > K - d

[0403] or, alternatively,

[0404] . if k < d

[0405]

[0406] = =b ifk > K - d In other words, for example, instead of using ak+ aj in the denominator, the maximum of ak (also designated as a_k) and aj (also designated as a J) can as well be used. The factor 2 in the nominator can then be discarded. Different calculations of mk(also designated as mjc), that follow a similar principle as the above mentioned are also possible.

[0407] These calculations of mk(n) may be an example for the calculation of the comodulation strength information 132, 246.

[0408] Short time averages are can, for example, be calculated as moving averages with a possible look-ahead of D samples and length L with

[0409] 1n+D

[0410] ai(n) = - ee(m)

[0411]

[0412] m=n+D-L+l

[0413] where the averaging length can depend on band index i and thus adjusted to the temporal resolution of the auditory system in the corresponding frequency region. This would closely resemble the use of inner products as in [3], Here, as an example, another averaging method using first order HR lowpass filters is proposed

[0414] ai n) = pat(n - 1) + (1 - p)e (n)

[0415] where the decay factor p, for example with 0 < p < 1, determines the time constant and also can depend on band index i. Different decay factors for averaging envelopes or products thereof may be used. This method has, among other things, the advantage of low complexity, which, for example, even allows sample-wise application. The resulting estimates of correlations can, for example, be regarded as inner products of exponentially weighted envelopes, for example giving the strongest weight to the latest samples. This

[0416] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX can lead to a fast adaptation to changes in the signal characteristics, so that usually no look-ahead is required.

[0417] The depicted calculations of ctj(n) may be examples for the calculation of a component of the average values 243 as in Fig. 2.

[0418] For example, the resulting comodulation strength finally can be converted to a comodulation coefficient 252 ck(n), for example within an interval [0, 1], for example by mapping and clipping:

[0419] ' 0 if ck(n) < 0

[0420] ,,. (n)

[0421] W)=_anrfCfe(n) = 1 if cKn) > 1

[0422] Mmax ^min

[0423]

[0424] T / c(n) else

[0425] with e.g. Mmin= 0.3 and Mmax= 1.

[0426] Other mappings of the comodulation strength information 246 to the comodulation coefficient 252 are also conceivable.

[0427] Integration of OMR in a Perceptual Model for Audio Coding

[0428] An example perceptual model based on a filter bank (e.g. the filter bank 215 from Fig. 2) which models the spectral decomposition in the inner ear is shown in Fig. 3. Its individual filters, e.g. filters 321, 322, 323, have bandpass characteristics, that, among other things, may for example resemble the frequency dependent sensitivity of locations on the basilar membrane with selected characteristic frequencies. Frequently, for example, HR bandpass filters are applied in the first stage, for example with center frequencies uniformly distributed on a perceptual frequency scale like, for example, ERB or BARK. They can be, for example, gammatone filters or filters with specifically designed frequency responses. As they can, for example, be operated at the full sampling rate of the input audio signal 310, the temporal resolution of the filter outputs can, for example, be adjusted to the resolution required for proper quantizer control in the audio coder. This can, for example, be achieved by averaging the magnitudes (in case of complex valued filters) or absolute values (in case of real valued filters) followed by downsampling.

[0429] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0430] 64 For example, this can be achieved with the average determinators 341, 342 and 343 and the downsamplers 353, 352, 351.

[0431] Since postmasking (e.g., post-masking) effects can decay more slowly than the impulse responses of the HR filters, an additional stage can model this, e.g., using an envelope detector (peak detector) with an appropriate decay factor dk

[0432] pk(m) = max(_ek(m), dfepfc(m - l))

[0433] For example, the decay factor may be the decay factor 256 or 591.

[0434] Its output can then, optionally, be fed to a final stage for mapping to the right level and / or the required frequency resolution for quantizer control, and (optionally) incorporating the influence of the threshold in quiet. The whole system then can deliver, for example, the quantizer step sizes, which then can be used in specified frequency ranges.

[0435] Once the comodulation coefficient 252 is estimated as described above or according to an embodiment of this invention, the masking release can be easily be implemented by controlling an attenuation factor 481 in each band by the corresponding comodulation coefficient 252 cfc(n) as shown in Fig. 4. Another example of the comodulation coefficient 252 is given in Fig. 2 and another example for the attenuation factor 261 is given in Fig. 2. The attenuation factor 481 and 261 gk(n) can for example be derived via a linear function such that there is no attenuation for cfc(n) = 0 and the strongest with a factor gminfor ck(n) = 1. For example:

[0436] 5 / c(n) = 1 - cfc(n)(l - gmin)

[0437] Corresponding non-linear mappings are also applicable here. The gain factor (e.g., the attenuation factor 261; 481) could also be applied before the postmasking (e.g., post¬ masking) modelling without much difference in the outcome, for example since the estimated comodulation itself is smoothened by the involved averaging operations.

[0438] However, informal listening experiments indicated that, preferably (but not necessarily), only mild atenuations should be used in order to avoid too strong reductions of the masking thresholds even for relatively stationary signals. On the other hand, for example, for low pitched pulsetrain like signals, a stronger attenuation would be favorable.

[0439] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX Therefore, a new approach was developed, which is inspired by the indications of a dip listening effect in the human auditory system, which, for example, benefits from (or even requires) faster decays of the masking thresholds than implemented standard postmasking (e.g., post-masking) decay time constants. Thus, in embodiments of this invention, the comodulation coefficient 252 ck(n) may now be used to control the postmasking (e.g., post-masking) decay time constant (varDecayTimeConst 254) and furthermore the decay factor 256 dkleading to the model structure shown in Fig. 5 (or in Fig 1, or in Fig. 2).

[0440] For example, for filter bank k, the scaling value dkmay, for example, be calculated the following:

[0441] varDecayTimeConst

[0442] = DECAY_TIMECONST_MAX - ck* (DECAY JTIMECONST_MAX - DECAY_TIMECONST_MIN)

[0443] and

[0444] 1000

[0445]

[0446] k' varDecayTimeConst * sampleRate Here, for example, the time constants could be the following:

[0447] DECAY JTMEC0NSTJ AX = 10 DECAY_TIMECONST_MIN = 2.4

[0448] For example, the decay factor may be the decay factor 256, 591

[0449] With the proposed procedure or embodiment, codecs, e.g., codecs with high temporal resolution, can adjust their quantizer step sizes accordingly. Codecs with lower temporal resolution, i.e. higher frame lengths, can, for example, use the minimum threshold within each frame, possibly after weighting to take into account the window shape of an inverse transform. Thus, for example, finer quantization would be resulting from dips in the masking thresholds rather than from high comodulation coefficients directly.

[0450] First Evaluation within an Audio Coding Framework

[0451] In order to get an impression on the capabilities of a comodulation controlled postmasking (e.g., post-masking), a listening test was carried out. It compared two models with and

[0452] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

[0453] C4 without the new approach within an MPEG-H based transform coding framework. Since, for example, this framework operates with variable bit rates, both were tuned by scaling their resulting masking thresholds in a way that they achieved equal average bit rates for a very long item concatenated from a multitude of test signals. Here this was set to 56 kbit / s as an example. Nevertheless, the bit rates for individual items used in the test can differ. Thus, the test can reveal which of the models better indicates the real time-varying thresholds. The results illustrated in Fig. 8 indicate an average improvement by applying comodulation controlled postmasking (e.g., post-masking).

[0454] Further conclusions and remarks

[0455] To conclude, for example, embodiments according to the invention create a perceptual model considering comodulation masking release by post-masking adaptation.

[0456] For example, an embodiment according to the invention is a perceptual model which uses the effect of “comodulation masking release”, in order to control the post-masking effect. For example, a basis is a model with a filterbank, e.g., HR filterbank. For example, the comodulation is computed from the normalized combination of covariance coefficients of the filterbank outputs. After the short time averaging with (or using) an HR filter of first order, the comodulation strength is computed based on this (or based from this; e.g. on the basis of the filterbank outputs or on the basis of a result of the short time averaging). This (e.g. comodulation strength) is used to control the post-masking, e.g. in an audio encoder or in a psychoacoustic model for audio quality evaluation.

[0457] Embodiments according to the invention are related to the field of psychoacoustic and / or to the field of filter banks and / or to the field of comodulation masking release and / or to the field of post-masking.

[0458] Embodiments according to the invention can be used in the technical field of audio coding. Embodiments according to the invention can, for example, be used in an audio encoder. Embodiments according to the invention can, for example, be used in proprietary coding methods.

[0459] Embodiments according to the invention can, for example, be used in psychoacoustic models for quality evaluation

[0460] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX Moreover, it should be noted that any of the embodiments and applications mentioned here may optionally be supplemented by any of the features, functionalities and details disclosed herein, both individually and taken in combination.

[0461] Also, any of the features, functionalities and details disclosed here may optionally be introduced into any other embodiment, both individually and taken in combination.

[0462] Embodiment according to Fig. 9

[0463] In the following a method 900 for providing a masking threshold information 901 will be described taking reference to Fig. 9.

[0464] The method 900 according to Fig. 9 is configured to obtain the masking threshold information 901 on the basis of an input audio signal 902.

[0465] The method comprises obtaining 910 envelope information on the basis of the input audio signal 902.

[0466] The method further comprises determining 920 comodulation strength information describing a comodulation between different frequency ranges of the input audio signal.

[0467] The method further comprises applying 940 a post-masking modeling to the envelope information or to a pre-processed version thereof, to obtain a post-masking processed envelope information, which describes a masking threshold, and adapting 930 a post¬ masking decay time of the post-masking modeling in dependence on the comodulation strength information.

[0468] The method may optionally be supplemented by any of the features, functionalities and details disclosed herein, both individually and in combination.

[0469] Embodiment according to Fig 10,

[0470] In the following an audio encoder will be described taking reference to Fig. 10.

[0471] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX The audio encoder 1000 according to Fig. 10 is configured to receive an input audio signal 1002 and provide, on the basis thereof, an encoded representation 1001 of the input audio signal 1002. For example, the audio encoder 1000 may comprise a quantization step size determinator 1010, that is configured to obtain a quantization step size information 1011 on the basis of the input audio signal 1002. The quantization step size determinator may, for example, be configured according to an embodiment of this invention. For example, the quantization step size determinator 1010 may correspond to the quantization step size determinator 500 according to Fig. 5.

[0472] The audio encoder may, for example, further comprise an optional preprocessing 1020 of the input audio signal 1002, that applies an audio preprocessing to the input audio signal 1002, to obtain a preprocessed version of the input audio signal 1021. The preprocessed version of the input audio signal 1021, or the input audio signal 1002 may further be processed by an optional time domain-to-spectral domain transformation / transformer 1030 to obtain a spectral representation of the input audio signal 1031. The spectral representation of the input audio signal 1031 may further be processed by an optional spectrum post-processing 1040 to obtain a post-processed version on the spectral representation of the input audio signal 1041.

[0473] Moreover, a quantization / quantizer 1050 may be configured to obtain quantized spectral values 1051 in dependence on the quantization step size information 1011 and on the basis of one of the following, depending on which optional steps have been performed: The post¬ processed version on the spectral representation of the input audio signal 1041, the spectral representation of the input audio signal 1031, the preprocessed version of the input audio signal 1021 or the input audio signal 1002. Optionally, a post-processing 1060 is applied to the quantized spectral values 1051 to obtain a post-processed version of the quantized spectral values 1061.

[0474] The encoded representation 1001 may then be derived from the post-processed version of the quantized spectral values 1061 or from the quantized spectral values 1051, for example using a spectral value encoding 1070. Optionally, the encoded representation 1001 may be supplemented by an encoded quantization information 1081 (e.g., encoded scale factors) that may be obtained by a quantization information encoding 1080.

[0475] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX However, it should be noted that the audio encoder may optionally be supplemented by any of the features, functionalities and details know in the art of audio encoding.

[0476] In addition, the audio encoder, or the quantization step size determinator 1010, may optionally be supplemented by any of the features, functionalities and details disclosed herein, both individually and taken in combination.

[0477] Implementation alternatives:

[0478] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.

[0479] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.

[0480] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0481] FH1IS24EM24 FHIIS24EM24-2024346428. DOCX Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.

[0482] 5

[0483] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0484] In other words, an embodiment of the inventive method is, therefore, a computer program 10 having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0485] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the 15 computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non- transitionary.

[0486] A further embodiment of the inventive method is, therefore, a data stream or a sequence of 20 signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.

[0487] A further embodiment comprises a processing means, for example a computer, or a 25 programmable logic device, configured to or adapted to perform one of the methods described herein.

[0488] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0489] 30

[0490] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 63

[0491] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0492] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0493] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0494] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software.

[0495] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0496] The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.

[0497] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0498] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 64

[0499] References

[0500] [1] B. Moore, An Introduction to the Psychology of Hearing: Fifth Edition (pp 100 - 105).

[0501] Emerald, 2008.

[0502] [2] J. Schnupp, I. Nelken, and A. J. King, Auditory Neuroscience: Making Sense of 5 Sound (pp 228 - 233). The MIT Press, 2011.

[0503] [3] A. J. S. Ferreira and D. Sinha, “A new broadcast quality low bit rate audio coding scheme utilizing novel bandwidth extension tools,” in Audio Engineering Society Convention 119, Oct 2005. [Online], Available: http: / / www.aes.org / e- I i b / b rowse. cf m?e I i b= 13341

[0504] 10 [4] S. Disch, S. van de Par, A. Niedermeier, E. Burdiel P'erez, A. Berasategui Ceberio, and B. Edler, “Improved psychoacoustic model for efficient perceptual audio codecs,” in Audio Engineering Society Convention 145, Oct 2018. [Online], Available: http: / / www.aes. org / elib / browse.cfm?elib=19755

[0505] 15

[0506] FHIIS24EM24 FHIIS24EM24-2024346428. DOCX

Claims

1. Claims1. A masking threshold determinator (100; 200; 301; 401; 501) for providing a masking threshold information (112; 280) on the basis of an input audio signal (110; 210; 310; 410; 510),3.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528) describing envelopes of different frequency ranges of the input audio signal (110; 210; 310; 410; 510);4.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to determine a comodulation strength information (132; 246) describing a comodulation between different frequency ranges of the input audio signal (110; 210; 310; 410; 510);5.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to apply a post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) to the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or to a pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), to obtain a post-masking-processed envelope information(142; 271; 366, 367, 368; 466, 467, 468; 566, 567, 568), which describes a masking threshold;6.wherein the masking threshold determinator (100; 200; 301; 401; 501 ) is configured to adapt a post-masking decay time of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) in dependence on the comodulation strength information (132; 246).

2. The masking threshold determinator (100; 200; 301; 401; 501) according to claim 1, wherein the masking threshold determinator (100; 200; 301; 401; 501 ) is configured to scale the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or a pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558)8.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 669., or the post-masking-processed envelope information (142; 271; 366, 367, 368; 466, 467, 468; 566, 567, 568), in dependence on the comodulation strength information (132; 246).

3. The masking threshold determinator (100; 200; 301; 401; 501 ) according to claim 1 or 2, 5 wherein the masking threshold determinator (100; 200; 301; 401; 501 ) is configured to scale the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or the pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), or the post-masking-processed envelope information (142; 271; 366, 367, 368; 466, 467, 468; 566, 567, 568), in dependence on the comodulation strength information (132; 10 246) using an attenuation factor (261),11.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to derive the attenuation factor (261) on the basis of a comodulation coefficient (152; 252)12.15 4. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 3,13.wherein the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) is configured to perform an envelope following with an adjustable decay time.14.20 5. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 4,15.wherein the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) is configured such that an output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 62, 463; 561, 562, 563) follows an increase of an input signal of the post-masking 25 modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) without delay, and such that the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) follows a reduction of the input signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) with a maximum step size which is an adjustable fraction of a of a previous output signal value of the post-masking modeling 30 (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) determined by a time constant value,16.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX17.CA 6718.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to adjust the time constant value in dependence on the comodulation strength information (132; 246).

6. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 5,20.wherein the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) is configured such that an output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) follows an increase of an input signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) without delay, and such that the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) follows a step-wise reduction of the input signal of the postmasking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) with an exponential decay,21.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to adjust a time constant of the exponential decay in dependence on the comodulation strength information (132; 246).

7. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 6,23.wherein the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) is configured to obtain the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) using a maximum function which determines a maximum value out of a current envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or a pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), and a scaled version of a previous output signal (270) of the postmasking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563),24.wherein the post masking modeling is configured to scale a previous output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) using a decay factor (256; 591) which is determined in dependence on the comodulation strength information (132; 246), to obtain the scaled version of the previous output signal (270) of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563).25.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX26. / 'X 8. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 7,27.wherein the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) is configured to obtain the output signal pfe(m)of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) according to28.Pk(m) = inax(e^m),dfePfc(m - 1))29.wherein e (m) is an envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or a pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558),30.wherein pk(m - l)is a previous value of the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563),31.wherein dkis a decay factor which is adjusted in dependence on the comodulation strength information (132; 246);32.wherein k is a frequency index;33.wherein m is a time index; and34.wherein max(...) is a maximum value operator.

9. The masking threshold determinator (100; 200; 301; 401; 501 ) according to one of claims 1 to 8,36.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain a plurality of sequences of frequency band values (211,212, 213; 311, 312, 313; 411, 412, 413; 511, 512, 513) representing an audio content in different frequency ranges; and37.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to derive the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528) from the sequences of frequency band values (211,212, 213; 311, 312, 313; 411, 412, 413; 511, 512, 513).38.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 10. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 9,39.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain the comodulation strength information (132; 246) using an evaluation of a covariance or of a correlation between temporal envelopes in two or more different frequency ranges.

11. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 10,41.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to determine the comodulation strength information (132; 246) using a computation of a correlation coefficient42.

43. according to44.wherein46.

47. Cki(n) = ek(n)'ei(tn') - afe(n) « / (n)48.wherein49.afc(n) = efc(n)50.wherein51.ai (n) = ei (n)52.wherein k is a frequency range index,53.wherein I is a frequency range index, with l = k - d or l = k + d,54.wherein d is a predetermined distance value,55.wherein n is a time index,56.wherein ek(n) is an envelope value in a frequency band having a frequency band index fc; wherein e{(n) is an envelope value in a frequency band having a frequency band index 1;57.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX wherein (...) is an averaging operation which performs a temporal averaging.

12. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 11,59.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to determine the comodulation strength information (132; 246) using a computation of a comodulation strength value mkaccording to60.G {dCiAn) + V? MT)61.mfe(n)63.

64. 2<u,(zs) + + a„(») or according to65.G(VCfc[(n) + yc^(n) )66.mfc(n)68.

69. max(ak(n), ai(n),au(n))70.or according to71.u if k < d72.nifc(n) = with j =74.

75. + aj(n) if k > K — d76.or according to77.if k < d78.mfe(n)80.

81. max(ak(n), aj(nf)WlJ [i if k > K - d82.wherein83.Ckt(n) = fife(n)ez(n) - ak(n)ai(n)84.wherein Ckuand Ckjare defined like Ckhwherein frequency index it or frequency index j takes the place of frequency index I,85.wherein86.akn) = ekin)'87.wherein88.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX wherein auand89.

90. are defined like abwherein frequency index u or frequency index j takes the place of frequency index I91.wherein k is a frequency range index,92.wherein I is a frequency range index, with k = k - d193.wherein u is a frequency range index, with u = k + d2,94.wherein and d2are predetermined distance values,95.wherein n is a time index,96.wherein ek(n) is an envelope value in a frequency band having a frequency band index c; wherein ez(n) is an envelope value in a frequency band having a frequency band index wherein eu(n) is an envelope value in a frequency band having a frequency band index u;97.wherein (...) is an averaging operation which performs a temporal averaging;98.wherein G is a normalization factor.

13. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 12,100.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to determine the comodulation strength information (132; 246) using a computation of a comodulation strength value,101.wherein the computation of the comodulation strength value comprises a computation of a quotient between102.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX - a sum of a plurality of covariance values describing a respective covariance or a respective correlation between temporal envelopes in two different frequency ranges, or a sum of a plurality of exponentiated versions of the covariance values, and103.- a sum, a weighted sum or a maximum of average values of sequences of envelope values in the different frequency ranges.

14. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 13,105.wherein the masking threshold determinator (100; 200; 301; 401; 501 ) is configured to map a comodulation strength value onto a comodulation coefficient (252) using a linear mapping and a limitation of a range of values.

15. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 14,107.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to use a smaller rate of samples in the post-masking modelling than in the determination of the comodulation strength information (132; 246).

16. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 15,109.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to apply an averaging and a downsampling to the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), in order to obtain a pre-processed version of the envelope information (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), wherein the masking threshold determinator is configured to apply the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) to the pre-processed version of the envelope information (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), to obtain the post-masking-processed envelope information (142; 271; 366, 367, 368; 466, 467, 468; 566, 567, 568); and110.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 73111.wherein the masking threshold determinator is configured to determine the comodulation strength information (132; 246) on the basis of the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), using a full temporal resolution of the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528).

17. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 16,113.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain a decay time constant value (254) in dependence on the comodulation strength information (132; 246) using a linear mapping.

18. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 17,115.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain a decay time constant value (254), which determines the post-masking decay time of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563), in dependence on a comodulation coefficient (252) using a linear mapping.

19. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 18,117.wherein the masking threshold determinator (100; 200; 301; 401; 501 ) is configured to scale a previous output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) using a decay factor (256; 591), to obtain a scaled version of the previous output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563), and118.to wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to use the scaled version of the previous output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) as a current output signal of the post-119.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) when slowing down a decay of the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563),120.wherein the masking threshold determinator is configured to determine the decay factor (256; 591) in dependence on a decay time constant value (254).

20. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 19,122.wherein the masking threshold determinator is configured to determine the decay factor (256; 591 ) in dependence on the decay time constant value (254), such that the decay factor (256; 591) increases with increasing decay time constant value (254).

21. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 20,124.wherein the masking threshold determinator is configured obtain the decay factor (256; 591 ) using a subtraction of a value representing a quotient between a sample period value and a decay time constant value (254) from a predetermined value.

22. The masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 21,126.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain the decay factor (256; 591) dkaccording to127.d — 1 — looo129. 130.,cvarDecayTimeConst*sampleRate’131.wherein varDecayTimeConst is a decay time constant value (254),132.wherein sampleRate is a value describing a sample rate of the output signal of the post¬ masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563).133.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 23. The masking threshold determinator according to one of the claims 1 to 22,134.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to obtain a decay factor (256; 591) dkaccording to135.. _1 NIOOO136.dk= 1.0 — — — — — — — -,137.varDecayTimeConst * sampleRate138.wherein varDecayTimeConst is obtained by139.varDecayTimeConst = DECAY_TIMECONST_MAX - ck' * (DECAY TIMECONST_MAX - DECAY_TIMECONST_MIN);140.wherein DECAY_TIMECONST_MAX and DECAY_TIMECONST_MAX are constant parameters for the decay time constant; wherein sampleRate is a value describing a sample rate of the output signal of the post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563);141.wherein ck' is a comodulation coefficient 252.

24. A quantization step size determinator (300; 400; 500) for providing a quantization step size information (311, 312; 411, 412; 511, 512) on the basis of an input audio signal (110; 210; 310; 410; 510),143.wherein the quantization step size determinator (300; 400; 500) comprises a masking threshold determinator (100; 200; 301; 401; 501) according to one of claims 1 to 23,144.wherein the masking threshold determinator (100; 200; 301; 401; 501) is configured to determine a masking threshold information (112; 280) on the basis of an input audio signal (110; 210; 310; 410; 510),145.wherein the quantization step size determinator (300; 400; 500) is configured to determine the quantization step size information (311, 312; 411, 412; 511, 512) using the masking threshold information (112; 280).146.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 25. The quantization step size determinator (300; 400; 500) according to claim 24, wherein the quantization step size determinator (300; 400; 500) is configured to determine the masking threshold information (112; 280) using a first frequency resolution, and wherein the quantization step size determinator (300; 400; 500) is configured to determine the quantization step size information (311, 312; 411, 412; 511, 512) using a second frequency resolution, which is different from the first frequency resolution.

26. An audio encoder (1000) for providing an encoded audio representation on the basis of one or more input audio signals,148.wherein the audio encoder (1000) comprises a quantization step size determinator (300; 400; 500) according to one of claims 24 to 25, wherein the quantization step size determinator (300; 400; 500) is configured to obtain the quantization step size information (311, 312; 411, 412; 511, 512) on the basis of the one or more input audio signals;149.wherein the audio encoder (1000) comprises a quantizer configured to quantize different frequency ranges of the one or more input audio signals using different effective quantization resolutions, to obtain quantized representations of the different frequency ranges of the one or more input audio signals; and150.wherein the audio encoder (1000) is configured to encode the quantized representations of the different frequency ranges of the one or more input audio signals, to obtain the encoded representation of the one or more input audio signals.

27. A method (900) for providing a masking threshold information (901) on the basis of an input audio signal (110; 210; 310; 410; 510),152.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX wherein the method (900) comprises obtaining envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528) describing envelopes of different frequency ranges of the input audio signal (110; 210; 310; 410; 510);153.wherein the method (900) comprises determining a comodulation strength information (132; 246) describing a comodulation between different frequency ranges of the input audio signal (110; 210; 310; 410; 510);154.wherein the method (900) comprises applying a post-masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) to the envelope information (122; 226, 227, 228; 326, 327, 328; 426, 427, 428; 526, 527, 528), or to a pre-processed version thereof (124; 224; 356, 357, 358; 456, 457, 458; 556, 557, 558), to obtain a post-masking-processed envelope information, which describes a masking threshold;155.wherein the method (900) comprises adapting a post-masking decay time of the post¬ masking modeling (140; 270; 361, 362, 363; 461, 462, 463; 561, 562, 563) in dependence on the comodulation strength information (132; 246).

28. A method for providing a quantization step size information (311, 312; 411, 412; 511, 512) on the basis of an input audio signal (110; 210; 310; 410; 510),157.wherein the method comprises providing a masking threshold information, according to claim 27,158.wherein the method comprises determining a masking threshold information on the basis of an input audio signal (110; 210; 310; 410; 510),159.wherein the method comprises determining the quantization step size information (311, 312; 411, 412; 511, 512) using the masking threshold information.160.FHIIS24EM24 FHIIS24EM24-2024346428. DOCX 29. A method for providing an encoded audio representation on the basis of one or more input audio signals (110; 210; 310; 410; 510),161.wherein the method comprises providing a quantization step size information (311, 312; 411, 412; 511, 512) according to claim 28,162.wherein the providing the quantization step size information (311, 312; 411, 412; 511, 512) comprises obtaining the quantization step size information (311, 312; 411, 412; 5 1, 512) on the basis of the one or more input audio signals (110; 210; 310; 410; 510);163.wherein the method comprises quantizing different frequency ranges of the one or more input audio signals (110; 210; 310; 410; 510) using different effective quantization resolutions, to obtain quantized representations of the different frequency ranges of the one or more input audio signals (110; 210; 310; 410; 510); and164.wherein the method comprises encoding the quantized representations of the different frequency ranges of the one or more input audio signals (110; 210; 310; 410; 510), to obtain the encoded audio representation of the one or more input audio signals (110; 210; 310; 410; 510).

30. A computer program for performing the method according to one of claims 27 to 29 when the computer program runs on a computer.166.FHIIS24EM24 FFniS24EM24-2024.346428. DOCX167.A.