Psychoacoustic Model for Audio Processing
The method addresses the challenge of inefficient bit allocation in audio encoding by using a masking model based on energy values and quiet-time auditory thresholds to improve bit allocation and audio quality.
Patent Information
- Application Number
- JP2022532788
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-12-03
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2040-12-03
AI Technical Summary
Existing audio encoding technologies face challenges in accurately calculating masking thresholds based on known characteristics of human hearing, leading to inefficient bit allocation across frequency bands and potential degradation in audio quality.
A method for processing audio signals that uses a masking model based on the energy value of the audio signal in a frequency band and the quiet-time auditory threshold for that band, calculating a sensitivity value and masking threshold to improve bit allocation and audio quality.
The proposed method more accurately captures the masking behavior of the human auditory system, leading to improved bit allocation and enhanced audio quality by reducing the likelihood of over- or under-assigning bits across frequency bands.
Smart Images

Figure 0007682884000003 
Figure 0007682884000004 
Figure 0007682884000005
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 943,903, filed on December 5, 2019, and European Patent Application No. 19213742.0, filed on December 5, 2019, the entire contents of both of which are incorporated herein by reference.
[0002] Technical Field The present disclosure relates to the field of audio processing, and in particular, to a method of processing an audio signal using a masking model based on the auditory threshold at rest in a frequency range of the audio signal and the measured energy value of the audio signal for the corresponding frequency range. The present disclosure further relates to an apparatus capable of performing the audio processing method.
Background Art
[0003] The human brain cannot memorize all audio signals at all different frequencies. Thus, when encoding audio, it is beneficial to remove signals at frequencies and levels that are imperceptible to the human auditory system. This is typically done by removing non - significant components from the audio signal. In the context of perceptual audio coders, there are two main ways for the encoder to increase compression efficiency. These are to remove the redundancy and non - significance of the signal. Redundant (predictable) signal components are typically removed in the encoder and restored in the decoder. Non - significant signal components are typically removed in the audio encoder by quantization and not restored by the audio decoder.
[0004] Typically, an encoder uses a psychoacoustic model, sometimes also called a perceptual model, to estimate the masking threshold for the audio spectrum. The masking threshold provides an estimate of the minimum just-noticeable distortion (JND) that is tolerable in each of multiple frequency bands of the audio spectrum. According to the critical bands of the human auditory system, the frequency bands are typically non-uniform in width. In a typical encoder, the masking threshold is input into a rate control loop that selects scale factors (and quantization noise levels) for each of multiple scale factor bands. The performance of a typical encoder depends on how close the masking threshold estimate is to the true JND noise level. If the masking threshold estimate exceeds the JND noise level, fewer bits than necessary to avoid audible distortion are allocated, and if the masking threshold estimate is below the JND noise level, more bits than necessary are allocated, potentially sacrificing neighboring frequency bands.
[0005] Typically, an encoder determines the masking threshold as follows: 1) Calculate the signal energy of the audio signal on a critical band frequency scale. 2) Estimate the frequency response of the signal after processing by the basilar membrane (also known as the excitation function) by convolving the critical band energy with a set of spreading functions. 3) In each critical band, adjust the excitation function of that band downward by the amount estimated to achieve the JND noise level.
[0006] Furthermore, models for determining the masking threshold often include heuristic rules that are developed experimentally and not directly based on known characteristics of human hearing.
[0007] Therefore, there is room for improvement in the technical field of calculating masking thresholds based on known characteristics of human hearing in order to improve bit allocation for various frequency bands of an audio signal.
Summary of the Invention
Problems to be Solved by the Invention
[0008] In view of the above, an object of the present disclosure is to overcome or mitigate at least some of the above problems. In particular, an object of the present disclosure is to provide a masking model based on the energy value of an audio signal in a frequency band and the quiet-time auditory threshold for that frequency band. Furthermore, an object of the present disclosure is to provide a masking model that reduces the complexity of audio encoding and improves the quality of the encoded audio based on the above. Further and / or alternative objects of the present disclosure will be apparent to the reader of the present disclosure.
Means for Solving the Problems
[0009] According to a first aspect, a method for processing an audio signal is provided. The audio signal includes audio data in a plurality of frequency bands, and the method includes: For each frequency band of the plurality of frequency bands: Determine an energy value for the audio data of that frequency band; Determine the quiet-time auditory threshold for that frequency band; Calculate a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold; Calculate a masking threshold for that frequency band using the sensitivity value and the energy value; Determining a bit allocation value for the frequency band using the energy value and the masking threshold.
[0010] Regarding the term "energy value", it should be understood that in the context of this specification, various approaches for calculating those energies can be used. The calculations can be based, for example, on the banded modified discrete cosine transform (MDCT), discrete Fourier transform (DFT), or complex MDCT (CMDCT). Note that for a certain frequency band, several energy values can be calculated and then combined in an appropriate way to form a single energy value for that frequency band. In this specification, "energy value" can refer to energy represented on either a linear or dB scale.
[0011] Regarding the term "frequency band", in the context of this specification, a frequency band should be understood as an interval within a frequency region that has a frequency range, delimited by a lower and an upper frequency, and has a certain frequency range. Note that the plurality of frequency bands of the audio signal to be encoded do not necessarily have the same width / range. For example, a relatively low frequency band may have a width of 100 - 200 Hz, and a relatively high frequency band may have a width of 3000 - 3500 Hz. Typically, the width of the frequency band increases with increasing frequency, and the frequency band between a relatively low frequency band and a relatively high frequency band may typically have a width anywhere in the range of 100 - 3000 Hz.
[0012] The term "sensitivity value" (SV) should be understood in the context of this specification as an approximation of the adjustment required in a given critical band to achieve JND distortion for a human listener with normal hearing. To account for the effect of masking across critical bands, the SV for each band can depend not only on the signal characteristics within that band but also on the signals in neighboring bands. The SV for each band is typically applied as an offset or adjustment to the excitation function, and then the quiet threshold is applied to derive the final masking threshold. All noise below the masking threshold is inaudible.
[0013] The SV for a particular frequency band can be calculated, for example, using the ratio, or difference, between the energy value of that frequency band and the quiet-time auditory threshold for that frequency band, or any other metric that compares the energy value and the quiet-time auditory threshold.
[0014] In a typical prior art encoder, the downward adjustment made to the excitation function in the critical frequency band is typically invariant with respect to the signal level, except for the application of the quiet-time threshold at the end. As a result, the estimated masking threshold may not fully correlate with the masking behavior of the human auditory system.
[0015] Thus, the representation of the adjustment for JND is typically level-independent. Such models are typically based on masking data for relatively large or relatively quiet signals. This approach can limit codec performance, for example, by underestimating the true JND threshold for low-level signal components and resulting in an excessive allocation of bits to frames containing relatively quiet signal passages. This problem occurs not only for variable bitrate encoders but also for encoders operating in a constant bitrate mode with a bit reservoir. Audio content characterized by very dynamic level changes, such as speech, is adversely affected.
[0016] In the present disclosure, by calculating the masking threshold based on both the SV and the energy value, the masking threshold can more accurately capture the observed masking behavior of the human auditory system, thereby enabling the delivery of higher quality audio signals.
[0017] Furthermore, when encoding an audio signal using a model that more faithfully captures the observed masking behavior of the human auditory system, the present method can more accurately estimate the number of bits necessary to meet a given quality target of providing an audio signal of a certain quality, thereby reducing the likelihood of over- or under-assigning bits. In embodiments where a certain bitrate is desired, this method can provide an audio signal of improved quality due to the improved bit-allocation strategy.
[0018] The present method can further provide a better match to subjectively measured masking data. Using the described audio encoding model, a single appropriate model can be achieved for all recorded speech levels or audio content. Advantageously, the model can facilitate the encoding of an audio signal of a certain quality regardless of the characteristics of the audio signal to be encoded. Some examples of audio signal characteristics are pitch, loudness, or duration, but note that there are many other characteristics to an audio signal.
[0019] General Description of Embodiments According to some embodiments, calculating the masking threshold involves applying a spreading function to one of the energy values for those frequency bands; or the transformed energy values of those frequency bands to determine an excitation value for the frequency band, and combining the sensitivity value with the excitation value.
[0020] The excitation function is thought to be the energy distribution along the basilar membrane of the inner ear. Thus, the excitation value is the value calculated from that function for a particular frequency band.
[0021] To mimic the processing of sound in the basilar membrane of the ear and smooth the predictability index over frequency, a diffusion function is applied to the energy value or the transformed energy value. For example, the diffusion function may be applied to the energy value converted to the loudness domain (i.e., raising the energy value to the power of about 0.25 to 0.3). In other embodiments, the diffusion function may be applied to the energy value raised to the power of 0.5 to 0.6. A diffusion function from ISO / IEC 11172-3:1993(E) may be used.
[0022] When the sensitivity value and the excitation value are defined in decibels (dB), the combining step may include calculating the masking threshold by subtracting the sensitivity value from the excitation value. In the intensity scale, the masking threshold is calculated as the quotient of the excitation value and the sensitivity value.
[0023] Optionally, the masking threshold is derived by threshold processing using the threshold at rest, for example, masking threshold = max(masking threshold, auditory threshold at rest).
[0024] According to some embodiments, the calculation of the masking threshold includes combining the energy value and the sensitivity value to determine an intermediate threshold and applying a diffusion function to the intermediate threshold to determine the masking threshold.
[0025] For example, the masking threshold may be determined as max(intermediate threshold, auditory threshold at rest).
[0026] According to some embodiments, the method further includes quantizing the audio samples of the audio data in the frequency band in response to the bit allocation value. Advantageously, the encoder can encode the audio with a certain quality or with improved audio quality at a certain bit rate. The encoder can further encode the quantized audio data in the frequency band into a bitstream.
[0027] The method described in this specification can also be used on the decoder side. According to some embodiments, the audio signal is an encoded bitstream that includes encoded energy values for frequency bands, and determining the energy value for the audio data in a frequency band includes decoding the encoded energy value from the encoded bitstream. On the decoder side, the determined bit allocation value can be used to extract the quantized audio samples of the audio data in that frequency band from the encoded bitstream. Advantageously, the bit allocation value for each frequency band of each audio frame does not need to be included in the bitstream and can instead be determined on the decoder side. In this way, the bit rate of the encoded bitstream can be reduced.
[0028] According to some embodiments, the method further includes dequantizing the quantized audio samples of the audio data in a frequency band and combining the dequantized audio samples of the audio data in each frequency band to generate a decoded audio signal.
[0029] According to some embodiments, determining the bit allocation value includes adjusting a masking threshold to achieve a bit allocation that satisfies a target bit rate for the audio signal. In this embodiment, if the number of bits required by the nominal masking threshold is more (or less) than the number of bits available to meet the bit rate requirement, the masking threshold can be adjusted to allocate more or fewer bits in order to use as many bits as possible without exceeding the target bit rate. For example, adjusting the masking threshold can include adjusting the masking threshold by adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied.
[0030] When determining and defining energy and auditory thresholds, different measurement values can be used. According to some embodiments, the energy value, the auditory threshold at rest, and the masking threshold are defined in decibels (dB). Since the decibel is a common metric for volume / energy, this provides simplicity to the model.
[0031] According to some embodiments, the method includes determining the plurality of frequency bands of the audio signal according to an Equivalent Rectangular Bandwidth (ERB) scale. The ERB scale provides an approximation of the bandwidth of the human auditory system using the convenient simplification of modeling the auditory filter as a rectangular band-pass filter. Advantageously, using ERB can be beneficial when encoding an audio signal according to the human auditory system.
[0032] According to some embodiments, the SV is defined in dB as a subtractive adjustment to the excitation function, and the step of determining the bit allocation value includes allocating more bits to the frequency band with a higher SV compared to the frequency band with a lower SV. Advantageously, a certain audio quality of the encoded audio signal can be achieved. The SV controls the displacement of the excitation function that gives the masking threshold after application of the threshold at rest. A positive sensitivity value lowers the masking threshold. A negative sensitivity value raises the masking threshold. As a result, an increasing sensitivity value corresponds to a lower masking threshold and thus to more bits being allocated. Thus, the sensitivity value for a frequency band can be considered to correspond to the sensitivity of the human auditory system to noise (encoding artifacts) in that frequency band of the audio signal.
[0033] According to some embodiments, the step of calculating the SV for the frequency band includes calculating a first SV using the sensation level, which is the difference on the dB scale between the energy value and the auditory threshold at rest.
[0034] The term "difference" should be understood in the context of this specification as the energy value (expressed in dB) minus the quiet-time auditory threshold (expressed in dB).
[0035] The term "sensation level" as used in this specification is defined as the level of a sound relative to its quiet-time threshold for an average listener. This term was introduced by Non-Patent Document 1.
Non-Patent Document 1
[0036] Note that the accurate determination of the SV can be achieved in various ways.
[0037] According to some embodiments, the step of calculating the first SV includes multiplying the sensation level by a first scalar. Advantageously, for the human auditory system, higher accuracy of the SV can be achieved with low complexity. By multiplying the first scalar by the function, the difference and the SV can be more easily mapped to each other so that they better correspond to the human auditory system.
[0038] The first scalar may be frequency-dependent or constant across all frequency bands.
[0039] According to some embodiments, the step of calculating the first SV includes adding a second scalar to the sensation level multiplied by the first scalar. The second scalar may be frequency-dependent or constant across all frequency bands.
[0040] Advantageously, for the human auditory system, higher accuracy of the SV can be achieved with lower complexity. By adding a second scalar to the function, the mapping between the difference and the SV can be easily changed to better correspond to the human auditory system.
[0041] According to some embodiments, the step of calculating the SV includes using a first SV as the SV for a frequency band.
[0042] According to some embodiments, the step of calculating the SV for a frequency band includes calculating a second SV using a perceptual level and weighting the first and second SVs based on at least one characteristic of the audio signal.
[0043] As an example, such characteristics of the audio signal can be bandwidth, tone-to-noise, nominal level, or power level in decibels (dB). However, it should be noted that the audio signal has many characteristics that can be used in this method. By weighting the first and second SVs, a masking threshold close to the JND can be obtained. Advantageously, a certain audio quality can be achieved regardless of the audio content of the audio signal.
[0044] In this embodiment, when using the model in which two SVs are calculated and weighted, a noisy, loud audio signal may result in an encoded high-quality audio signal. A soft audio signal that is not so noisy may also be able to provide the same level of quality for the encoded audio signal when using the same model. Conversely, in the prior art dual-mode encoder, for an audio signal that cannot be reliably classified, such as a dialog mixed with applause, the encoder has to select either, for example, the applause mode or the default mode, and the mode selected by the encoder may not be optimal. Alternatively, the encoder may choose to encode different segments of the signal in different modes, which may result in audible switching artifacts that degrade the quality of the encoded signal. In a single-mode model in which the first and second SVs are calculated and weighted, there is no need to determine which part of the audio signal should be encoded in different modes. Furthermore, the single-mode model is suitable for a variety of audio signals that do not depend on the audio content (speech, music, etc.) of the audio signal. By providing a single model, there is no need to classify audio samples to determine which mode needs to be applied for encoding. Furthermore, the problem of selecting a mode for signals at the boundary between modes is alleviated, and the degradation of audio quality due to the encoder making a non-optimal mode selection is avoided.
[0045] Note that in some embodiments, additional SVs may be defined and can be included (by weighting more than two SVs) when calculating the final SV. These different SVs may be weighted based on the transient characteristics of the audio signal.
[0046] According to some embodiments, the step of calculating the second SV for a frequency band includes multiplying a third scalar, different from the first scalar, by the perceived level. The third scalar may be frequency-dependent or may be constant across all frequency bands.
[0047] By multiplying the third scalar to the function, it is possible to improve the accuracy of the SV regarding the auditory system for audio signals having different characteristics, such as an audio signal with a high degree of noise-like or tone-like characteristics. The third scalar may allow the SV to be mapped in order to enable a generalized model for audio coding. The SV is defined as a function of the perceived level as understood from the above and further described below. By mapping the SV according to different characteristics of the audio signal, the relationship between the auditory threshold and the energy value can remain the same, while the slope of the SV and the perceived level may vary according to different characteristics of the audio signal. Advantageously, bits can be allocated to provide a high-quality, encoded audio signal.
[0048] According to some embodiments, the step of calculating the second SV includes adding a fourth scalar to the perceived level multiplied by the third scalar, where the fourth scalar is different from the second scalar. It should be noted that the fourth scalar may be assigned various values to improve the accuracy of the SV regarding the human auditory system. The fourth scalar may be frequency-dependent or may be constant across all frequency bands.
[0049] According to some embodiments, the step of weighting the first and second SVs based on at least one characteristic of the audio signal includes calculating a value representing a weight, the value being in the range between 0 and 1, and the step of calculating the SV for a frequency band includes multiplying the value by one of the first and second SVs and multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums together to form the SV for that frequency band.
[0050] In other words, the first SV and the second SV are mixed as a linear combination using weights whose sum is 1, and the weights depend on the at least one characteristic of the audio signal.
[0051] By calculating the overall SV by weighting the first and second SVs, the mapping between the perceptual level and the SV can be adapted to reflect the masking characteristics of different audio signal types. Advantageously, a model with low complexity and flexibility can be achieved for determining the bit allocation value of the frequency band, which provides a high-quality encoded audio signal.
[0052] According to some embodiments, the at least one characteristic defines the estimated tonality of that frequency band of the audio signal.
[0053] Note that there are many different ways to estimate the tonality of an audio signal. Tonality represents the relationship between tonal characteristics (such as notes, chords, keys, pitch, etc.) in an audio signal. Advantageously, using the estimated tonality of that frequency band of the audio signal as a characteristic for weighting the first and second SVs can improve the accuracy of the SV with respect to the human auditory system. Furthermore, using tonality may improve the subjective audio quality.
[0054] According to some embodiments, the at least one characteristic defines an estimated level of noise in that frequency band of the audio signal. Advantageously, the noise of the audio signal may be masked in relation to the human auditory system in order to achieve a high-quality encoded audio signal.
[0055] According to some embodiments, the estimated tonality is calculated using an adaptive prediction of frequency coefficients calculated from that frequency band of the audio signal. In some embodiments, the same set of frequency coefficients is used for both the calculation of the masking threshold and the estimation of tonality. In other embodiments, the estimation of tonality is performed using a separate complex-valued filter bank. Note that any set of frequency coefficients is possible depending on the desired accuracy of the estimated tonality and the available computational resources. For example, using only the real MDCT coefficients is computationally less expensive than using the CMDCT coefficients but is less accurate. By obtaining an accurate estimated tonality, the subjective audio quality can be further improved.
[0056] According to some embodiments, linear prediction coding LPC is adaptively applied to the MDCT coefficients. The application is based on the frequency band of the audio signal from which the MDCT coefficients are calculated. In order to achieve a more accurate tonality estimate of the audio signal, LPC may be used instead of a fixed prediction.
[0057] Note that the LPC analysis window may have different lengths. By varying the analysis window length, a desired variable time-frequency framework can be flexibly realized. According to some embodiments, the LPC analysis window length is varied as a function of the frequency band. In some embodiments, a relatively long LPC analysis window is used for relatively lower frequency bands.
[0058] According to some embodiments, the prediction order of the LPC can be varied as a function of the frequency band. As an example, the prediction order of the LPC can be selected such that the distinction between pure noise input and signals with tonal components (such as harpsichord, speech, etc.) is maximized.
[0059] Note that the frequency range of the audio signal may be encoded into different ranges.
[0060] According to some embodiments, the frequency range of the audio signal is between 200 and 7000 Hz.
[0061] According to some embodiments, the step of determining the quiet-time auditory threshold for the frequency band includes using a predefined table that defines the auditory threshold for at least some frequencies. The predefined table can be pre-stored in the encoder that executes this method, thereby allowing the predefined table to be updated without affecting the decoder compatibility. Advantageously, the complexity of providing a high-quality encoded audio signal can be reduced.
[0062] According to some embodiments, before quantizing the audio samples of the audio data in the frequency band, the dynamic range of the audio signal is reduced using a companding algorithm. By companding the audio signal in the encoder and applying complementary expansion in the decoder, the encoding method can provide a higher-quality decoded audio signal. Companding the audio signal allows for fewer bits to be encoded while maintaining high audio quality.
[0063] According to some embodiments, the method includes defining a spreading function for a frequency band depending on a perceived level such that an effect of the spreading function in a frequency band having a relatively high perceived level is greater compared to an effect of the spreading function in a frequency band having a relatively low perceived level.
[0064] According to a second aspect, at least one of the above objectives is achieved by a device, which: a receiving component configured to receive an audio signal, the audio signal including audio data in a plurality of frequency bands, and an analysis component; an analysis component configured to determine a plurality of frequency bands of the audio signal, wherein the analysis component, for each frequency band of the plurality of frequency bands: determines an energy value for the audio data of that frequency band; determines a quiet-time auditory threshold for that frequency band; calculates a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold; calculates a masking threshold for that frequency band using the sensitivity value and the energy value; is configured to determine a bit allocation value for the frequency band using the energy value and the masking threshold.
[0065] According to some embodiments, the analysis component: applies a spreading function to either the energy values of those frequency bands; or the converted energy values for those frequency bands to determine an excitation value for that frequency band, and is configured to calculate the masking threshold by combining the sensitivity value with the excitation value.
[0066] According to some embodiments, the analysis component is configured to calculate the masking threshold by combining the energy value and the sensitivity value to determine an intermediate threshold and applying a diffusion function to the intermediate threshold to determine the masking threshold.
[0067] According to some embodiments, the apparatus is an encoder, and the apparatus further has an encoding component configured to quantize audio samples of audio data in a frequency band in response to the bit allocation value.
[0068] According to some embodiments, the encoding component is further configured to encode the quantized audio data in the frequency band into a bitstream.
[0069] According to some embodiments, the apparatus is a decoder, the audio signal is an encoded bitstream including an encoded energy value for a frequency band, the apparatus further has a decoding component configured to decode the encoded energy value from the encoded bitstream, and the analysis component uses the decoded energy value when determining the energy value.
[0070] According to some embodiments, the decoding component is further configured to extract quantized audio samples of audio data in a frequency band from the encoded bitstream in response to the bit allocation value.
[0071] According to some embodiments, the decoding component is further configured to dequantize the quantized audio samples of audio data in a frequency band and combine the dequantized audio samples of audio data in each frequency band to generate a decoded audio signal.
[0072] According to some embodiments, when determining the bit allocation value, the analysis component is configured to adjust the masking threshold so as to achieve a bit allocation that satisfies a target bit rate for the audio signal.
[0073] According to some embodiments, when adjusting the masking threshold, the analysis component is configured to adjust the masking threshold by adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied.
[0074] According to some embodiments, the analysis component is configured to define the energy value, the auditory threshold at rest, and the masking threshold in decibels (dB).
[0075] According to some embodiments, the analysis component is configured to determine the plurality of frequency bands of the audio signal according to an Equivalent Rectangular Bandwidth (ERB) scale.
[0076] According to some embodiments, the SV is defined in dB as a subtractive adjustment to the excitation function, and the analysis component is configured to determine the bit allocation value by allocating more bits to frequency bands having a higher SV than to frequency bands having a lower SV.
[0077] According to some embodiments, the analysis component is configured to calculate the SV for a frequency band by calculating a first SV using a sensation level, where the sensation level is the difference between the energy value and the auditory threshold at rest.
[0078] According to some embodiments, the analysis component is configured to calculate the first SV by multiplying the first scalar by the sensory level.
[0079] According to some embodiments, the first scalar is frequency-dependent.
[0080] According to some embodiments, the first scalar is constant across all frequency bands.
[0081] According to some embodiments, the analysis component is configured to calculate the first SV by adding a second scalar to the sensory level multiplied by the first scalar.
[0082] According to some embodiments, the analysis component is configured to calculate the SV by using the first SV as the SV for its frequency band.
[0083] According to some embodiments, the analysis component is configured to calculate a second SV using the sensory level, and calculate the SV for its frequency band by weighting the first and second SVs based on at least one characteristic of the audio signal.
[0084] According to some embodiments, the analysis component is configured to calculate the second SV for its frequency band by multiplying a third scalar, different from the first scalar, by the sensory level.
[0085] According to some embodiments, the analysis component is configured to calculate the second SV by adding a fourth scalar, different from the second scalar, to the sensory level multiplied by the third scalar.
[0086] According to some embodiments, the analysis component weights the first and second SVs based on at least one characteristic of the audio signal, calculates a value representing the weight, the value being in the range between 0 and 1, multiplies one of the first and second SVs by the value, multiplies the other of the first or second SVs by 1 minus the value, adds the two resulting sums together to form an SV for that frequency band.
[0087] According to some embodiments, the at least one characteristic defines an estimated tonality of that frequency band of the audio signal.
[0088] According to some embodiments, the at least one characteristic defines an estimated level of noise in that frequency band of the audio signal.
[0089] According to some embodiments, the analysis component is configured to calculate the estimated tonality using an adaptive prediction of frequency coefficients calculated from that frequency band of the audio signal.
[0090] According to some embodiments, the analysis component is configured to adaptively apply LPC to MDCT coefficients based on a certain frequency band of the audio signal. The frequency band is the band from which the MDCT coefficients are calculated.
[0091] According to some embodiments, the LPC analysis window length is varied as a function of the frequency band.
[0092] According to some embodiments, a relatively long LPC analysis window is used for relatively lower frequency bands.
[0093] According to some embodiments, the prediction order of LPC is varied as a function of the frequency band.
[0094] According to some embodiments, the frequency range of the audio signal is between 200 and 7000 Hz.
[0095] According to some embodiments, the apparatus further comprises a memory that stores a table defining the hearing threshold at rest for at least some frequencies, and the analysis component is configured to determine the hearing threshold at rest for a frequency band by using the predefined table.
[0096] According to some embodiments, the apparatus further has a companding component configured to reduce the dynamic range of the audio signal using a companding algorithm before quantizing the audio samples of the audio data in various frequency bands.
[0097] According to some embodiments, the analysis component is configured to define a spreading function for the frequency band depending on the perceived level such that the effect of the spreading function in a frequency band having a relatively high perceived level is greater compared to the effect of the spreading function in a frequency band having a relatively low perceived level.
[0098] According to some embodiments, the apparatus is implemented as a real-time two-way communication device.
[0099] The second aspect may generally have the same advantages as the first aspect.
[0100] According to a third aspect, a method for estimating the tonality of an input signal is provided: applying a filter bank to achieve a set of frequency coefficients; calculating the estimated tonality using an adaptive prediction of the frequency coefficients.
[0101] According to some embodiments, the step of calculating the estimated tonality includes applying adaptive linear prediction LPC to the frequency coefficients based on the frequency band of the audio signal from which the frequency coefficients are calculated.
[0102] According to some embodiments, the LPC analysis window length is varied as a function of the frequency band.
[0103] According to some embodiments, a relatively long LPC analysis window is used for relatively low frequency bands.
[0104] According to some embodiments, the prediction order of the LPC is varied as a function of the frequency band.
[0105] According to some embodiments, the filter bank includes one of a 128-band complex MDCT or DFT filter bank and a 64-band complex quadrature mirror filter CQMF filter bank.
[0106] According to some embodiments, the LPC analysis window is an asymmetric Hamming window.
[0107] According to some embodiments, the method: includes weighting the predictability indicators from the adaptive prediction according to the relative perceptual importance of each predictability indicator.
[0108] According to some embodiments, the step of weighting the predictability indicators included within each time-frequency tile includes one of weighting based on the energy or loudness of the input signal.
[0109] According to some embodiments, the method: further includes combining the predictability indicators from the adaptive prediction of the frequency coefficients to match the time and frequency resolution of the filter bank.
[0110] Furthermore, it should be noted that, unless explicitly stated otherwise, the present disclosure relates to any possible combination of features.
Brief Description of the Drawings
[0111] The above, as well as additional objects, features, and advantages of the present disclosure, will be better understood through the following illustrative and non-limiting detailed description of embodiments of the present disclosure with reference to the accompanying drawings. Here, the same reference numbers are used for similar elements.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Best Mode for Carrying Out the Invention
[0112] Here, the present disclosure will be more fully described with reference to the accompanying drawings in which embodiments of the present disclosure are shown. The systems and devices disclosed herein are described in terms of their operations.
[0113] In the following, known audio formats are used as a context for exemplifying the present disclosure. However, it should be noted that the scope of the present disclosure is not limited to this known format, and the various embodiments described herein can be used for any suitable audio format.
[0114] For exemplary formats, there are currently two commonly used modes for encoding audio. Selecting which mode is most suitable for an audio signal can be a complex decision, and if a mode that is not suitable for the audio signal is selected, the quality of the encoded audio signal may degrade. The two typical modes are default and applause. The current modes are separate, and in both modes, the encoder estimates the masking threshold from the SV that is invariant to the signal energy estimate and the signal level, with the exception of applying the threshold at quiet times at the end. The default mode further applies legacy functions inherited from the MPEG layer III encoder, but there is not sufficient perceptual justification for this function. Further, the masking threshold is input to a rate control loop that selects scale factors (and quantization levels) for each of a plurality of scale factor bands. Thereby, the performance depends on how close the masking threshold estimate is to the true JND noise level.
[0115] In most prior art models, prior to the application of the quiet-time threshold, the required SMR representation for JND is level-independent. Such models typically rely on masking data for relatively loud or relatively quiet signals, but neither is adaptive. This approach can limit codec performance in one example by underestimating the true JND threshold for low-level signal components, resulting in an excessive allocation of bits to frames containing relatively quiet signal messages. This problem occurs not only for variable bitrate encoders, but also for encoders operating in constant bitrate mode with a bit reservoir. Audio content characterized by very dynamic level changes (e.g., speech) will be adversely affected.
[0116] A common problem with prior art models is that they generate masking thresholds lower than necessary, which also leads to an excessive allocation of bits within the frequency band. Thus, this reduces the number of bits available for other bands, thereby degrading the quality of the encoded audio signal.
[0117] The present disclosure aims to avoid some of the above problems by providing a single model that performs equally well or better than prior art dual-mode or single models for most audio content by estimating a more accurate SMR.
[0118] Subjective listening tests using monaural content have shown that the new encoder performs better than the current encoder for speech content. Further, this new encoder is significantly more effective in variable bitrate applications that provide a constant audio quality by allocating only the number of bits necessary for the encoder to meet a predefined quality goal.
[0119] In one experiment, to quantify the advantages of level-dependent masking, a first subjective listening test was conducted using three encoders and a set of diverse audio test items. Of the three encoders, one operates in the default mode, one operates in the clap mode, and the other operates using level-dependent masking. The encoder using level-dependent masking resulted in an average subjective quality improvement of 3 points and 14 points, respectively, compared to the default and clap mode encoders. More importantly, level-dependent masking improved two speech items by an average of 8 points compared to the default encoder.
[0120] FIG. 7 shows, by way of example, the analysis component 700. As will be further described below in connection with FIGS. 8 and 9, the analysis component may be implemented in the encoder 800 or the decoder 900. In other embodiments, the analysis component is implemented in a separate device and is connected, for example, to the encoder or the decoder.
[0121] The analysis component 700 has a circuit configured to execute a method of processing an audio signal to determine bit allocation values for various frequency bands of the audio signal. The circuit may include one or more processors.
[0122] The analysis component 700 is configured to perform various actions exemplified below.
[0123] The analysis component 700 is configured to determine S02 a plurality of frequency bands of the input audio signal. Each of the plurality of frequency bands includes a frequency range. It should be noted that each of the plurality of frequency bands of the audio signal to be encoded does not necessarily have the same width / range. In one example, the first relatively low frequency band may have a range of 100 to 200 Hz, and another relatively high frequency band may have a range of 3000 to 3500 Hz. In some embodiments, the frequency range of the audio signal may be 200 to 7000 Hz. Further, it should be noted that there are many different frequency ranges in the audio signal that may extend up to frequencies higher than 7000 Hz and / or lower than 200 Hz. As will be understood, there are various ways to determine the frequency bands for an audio signal. In some embodiments, the analysis component 700 is configured to determine S02 the frequency bands according to the equivalent rectangular bandwidth ERB scale. The ERB scale provides an approximation of the bandwidth of the filters of the human auditory system. Further, using the ERB scale provides the simplification of modeling the filters as rectangular band-pass filters.
[0124] The analysis component is further configured to determine S18 a bit allocation value for each frequency band using the following analysis of the audio data of each frequency band.
[0125] The analysis component 700 determines S04 an energy value for the audio data of the frequency band. The energy value may be, for example, the banded MDCT energy.
[0126] Furthermore, the analysis component 700 determines S06 the quiet-time auditory threshold for the frequency band. In some embodiments, the analysis component 700 comprises or is connected to a memory component. The memory component stores a table defining the quiet-time auditory threshold for at least some frequencies. It should be noted that such a memory component can store different information. In other words, determining S06 the quiet-time auditory threshold for the frequency band can include using a predefined table that defines the auditory threshold for at least some frequencies. In some embodiments, the predefined table that defines the auditory threshold can be replaced, allowing the encoder to be improved without affecting the decoder's compatibility.
[0127] Using the energy value and the quiet-time auditory threshold, a sensitivity value (SV) can be calculated S08. It should be understood that the SV can be calculated S08 in various ways using the energy value and the quiet-time auditory threshold. The SV can be calculated S08, for example, using the ratio, or difference, between the energy value and the quiet-time auditory threshold, or any other metric that compares the energy value and the quiet-time auditory threshold. The sensitivity value should be understood as a quantity defined, for example, in dB.
[0128] In one embodiment, the first SV is calculated S10 using the difference between the energy value and the hearing threshold in quiet (also referred to as the "sensation level" in the present disclosure). Optionally, the first SV may be calculated S10 by multiplying the sensation level by a first scalar. In some embodiments, the first SV may be calculated S10 by adding a second scalar to the difference multiplied by the first scalar. Thus, in this embodiment, the first SV for a certain frequency band is calculated as alpha*(band energy - hthresh) + beta. Here, alpha is the first scalar, beta is the second scalar, band energy is the energy value of the audio signal in that frequency band, and hthresh is the threshold in quiet for that frequency band. In some embodiments, the second scalar is not included in the calculation of the SV to reduce complexity.
[0129] The degree to which the first SV varies with the difference between the energy value and the threshold in quiet for different frequency bands is determined by examining various measured masking data. It should be noted that the following measurements and figures, described in connection with FIGS. 1 - 3, are provided as examples using the difference between the energy value and the hearing threshold in quiet, i.e., the sensation level, in each frequency band. However, those skilled in the art will understand that other data may be obtained from experiments if other ways of calculating the SV are used, for example, using the ratio between the energy and the hearing threshold for each frequency band.
[0130] Figure 1 shows, as an example, the measured masking data for a 200 Hz tone at different sound pressure levels SPL. The masking thresholds 104, 106, 108, 110, 112 are presented in relation to the auditory threshold 102 (thick line) at rest. The masking threshold 1 104 relates to the 200 Hz tone at 60 dB SPL. The masking threshold 2 106 relates to the 200 Hz tone at 80 dB SPL. The masking threshold 3 108 relates to the 200 Hz tone at 90 dB SPL. The masking threshold 4 110 relates to the 200 Hz tone at 100 dB SPL. The masking threshold 5 112 relates to the 200 Hz tone at 105 dB SPL. As can be seen from Figure 1, the difference between the level of the tone - masker (i.e., 60, 80, 90, 100, 105 dB; although not specifically shown in Figure 1, it can be easily found by following the markings on the vertical axis from each sound level to the 200 Hz marking on the horizontal axis) and the masking threshold at 200 Hz increases with an increase in the intensity of the sound of the tone - masker. For a 60 dB tone - masker, the difference is approximately 18 dB (the tone - masker is 60 dB, the masking threshold 104 is 42 dB), and for a 105 dB tone - masker, the difference is approximately 32 dB (the tone - masker is 105 dB, the masking threshold 112 is 73 dB).
[0131] Figure 2 shows a similar pattern for a 500 Hz tone - masker. In Figure 2, the masking thresholds 204, 206, 208, 210 (the same levels as in Figure 2, except that the 105 dB tone - masker is not shown in Figure 2) for masking data at different sound pressure levels SPL are presented in relation to the auditory threshold 102 (thick line) at rest.
[0132] In one example, the measured masking data (i.e., such as those illustrated in FIGS. 1 and 2) can be used to derive the sensitivity value SV versus the sensation level when a tonal or sine wave signal is considered. As will be appreciated, there are parameters other than those illustrated in FIGS. 1 - 2 that can be used to derive the required SV. In this example, masking for coding artifacts within the auditory critical band at the sine wave masker frequency was considered.
[0133] FIG. 3 shows, as an example, a model of the SMR as the combination of tonal - masking narrow - band noise for different tonal sensation levels and frequencies. The SMR is presented using a linear model 302 as a function of the sensation level. The linear model 302 is a combination of the measured SMR curves 1 - 6, which are referred to as 304, 306, 310, 312, 314, 316 in FIG. 3. Curve 1 304 shows the measured SMR values for a signal at a frequency of 200 Hz. Curve 2 306 shows the measured SMR values for a signal at a frequency of 500 Hz. Curve 3 310 shows the measured SMR values for another signal at a frequency of 500 Hz. Curve 4 312 shows the measured SMR values for a signal at a frequency of 1000 Hz. Curve 5 314 shows the measured SMR values for a signal at a frequency of 2000 Hz. Curve 6 316 shows the measured SMR values for a signal at a frequency of 5000 Hz.
[0134] FIG. 3 shows for this example that a slope of SMR versus sensation level of 0.35 dB*(masker level relative to the quiet threshold at that frequency)+3 dB can be a reasonable approximation for the frequency range 200 - 4000 Hz. However, it should be noted that the linear model 302 in FIG. 3 can be a reasonable approximation for other frequency ranges. In this example, the decibel offset for the required SMR varies up to 10 dB at intermediate levels, but appears to converge at high and low levels.
[0135] In some embodiments, the quiet threshold may be modified by setting the threshold for all bands below 4 kHz to a global minimum threshold. The quiet threshold should be set to the minimum value within each band during encoding. For example, in a transform codec with adaptive block switching, the lowest frequency band of the shortest transform block may be 750 Hz wide. As can be seen from FIGS. 1-2, the level (dB) of the auditory threshold at quiet rapidly decreases from 20-750 Hz. And the threshold for this entire band can be set to the actual quiet threshold at 750 Hz. The same step is applied to all other bands in the shortest block. And these values are interpolated to obtain the quiet threshold for all other transform block lengths. This approach ensures that the quiet threshold is at a consistent level for all block lengths and avoids unwanted quantization noise modulation artifacts when the codec switches the transform length. An alternative, simpler approach is to set the threshold for all bands below 4 kHz to the global minimum threshold. Using this adjusted quiet threshold will result in other values for the first and / or second scalars, as will be understood by those skilled in the art.
[0136] Note that the quiet threshold in FIGS. 1-2 is placed conservatively 20 dB below what would be expected under normal assumptions of a peak playback level of 105 dB SPL. In one embodiment, the threshold is set based on a peak playback level of 115 dB. This provides some robustness when playing back decoded audio at a level different from the assumed level, particularly for variable bitrate applications.
[0137] The model of FIG. 3 is derived by averaging the results of tone masking narrowband noise experiments for various frequencies. Signal components with higher sensation levels receive higher SVs. In one example, the SV increases by 1 dB for every 3 dB increase in the band energy above the quiet-time auditory threshold. The level-dependent SV model SV(j) in FIG. 3 for frequency band j is expressed as follows: SV(j)=max(0,0.35*(Eb(j)-Q(j))+3) where Eb(j) and Q(j) (expressed in dB in this example) are the banded MDCT energy and the quiet-time threshold, respectively.
[0138] As will be appreciated, it will be apparent that by changing the configuration of the analysis components, the scalars presented in the above equations can be changed. Those scalars may be modified to adjust the calculation of the SV to better fit some audio signals. The first scalar may be, for example, in the range of 0.2 to 0.5. The second scalar may be, for example, in the range of 2.5 to 3.5.
[0139] In the model of FIG. 3, the first scalar and the second scalar are constant across all frequency bands. However, in other embodiments, the first and / or second scalar is frequency-dependent.
[0140] In one embodiment, as shown in FIG. 3, a first SV is calculated S10 and used as the SV for that frequency band.
[0141] In one embodiment, the linear model 302 in FIG. 3 is extended to more accurately estimate the SV for an input signal with a high level of noise. Some examples of input signals with a high noise level can be clapping, or rain, or the fricative sound of speech. However, as will be appreciated, there are many different signals with a high noise level. As an example, when using the same methodology as in the case of tone-masking-noise, the SV can be calculated to be more accurate for some signals.
[0142] In a second embodiment, it is proposed that at a high noise level, the same linear relationship exists between the SV and the sensory level, but the slope is different. The slope of the best-fit line is approximately half that in the case of tone-masking-noise. This correspondence has been verified using experiments similar to those shown in FIGS. 1 - 3, but for a noise masker instead of a tone masker. Thus, a generalized model can be realized by adapting the SV rule depending on the input signal characteristics.
[0143] Thus, in one embodiment, an optional second SV is calculated S12 and combined with the first SV using an optionally fixed or adaptively weighted combination to define the final SV. In these embodiments, the analysis component 700 calculates S12 the second SV using the difference (sensory level) between the energy determined S04 and the quiet-time auditory threshold determined S06 when calculating S08 the SV for the frequency band, and is further configured to weight the first SV and the second SV based on at least one determined S14 characteristic of the input audio signal. As will be appreciated, any suitable characteristic of the audio signal can be used in the calculation S08 of the SV. In one embodiment, the at least one characteristic is the estimated tonality of the signal. Alternatively, in one embodiment, the at least one characteristic is the estimated noise level for the signal.
[0144] In one embodiment, the estimated tonality is calculated using an adaptive prediction of frequency coefficients computed from the frequency band of the audio signal. Embodiments for estimating the tonality of an audio signal are described below.
[0145] As will be appreciated, it is possible to use any set of frequency coefficients. As an example, one prior art method is based on a second order fixed prediction over time of the absolute value and phase of the DFT (Non-Patent Document 2). [Non-Patent Document 2] According to ISO / IEC 11172-3:1993(E), "Information technology - Coding of moving pictures and associated audio for digital storage media at up to about 1.5 Mbit / s - Part 3: Audio", overlapping DFTs of lengths 512 and 128 (i.e., the number of complex DFT coefficients) are computed in parallel to enable different time / frequency resolution trade-offs for different frequencies. Analysis component 700 can use adaptive linear prediction of complex MDCT (CMDCT) coefficients by generalizing prior art methods. In some embodiments, linear predictive coding (LPC) can be adaptively applied to the MDCT coefficients based on the frequency band of the audio signal for which the MDCT coefficients are computed. Adaptive linear prediction allows rapidly evolving midrange harmonics in voiced speech and music to produce higher tonal estimates than fixed prediction. Additionally, the desired variable time / frequency framework can be flexibly implemented without the need for a parallel CMDCT filter bank by varying the LPC analysis window length and / or the prediction order as a function of frequency. In other words, the LPC analysis window length can be varied as a function of the frequency band. Further, the prediction order of the LPC can also be varied as a function of the frequency band. Optimal LPC analysis parameters may be selected offline for each frequency band by maximizing the difference in the average prediction gain for a challenging signal and independent and identically distributed (IID) Gaussian noise. Examples of challenging signals may be speech or harpsichord. However, it should be understood that there are many different signals that can be classified as challenging. The longest LPC analysis window is typically used at low frequencies, and progressively shorter windows are used at higher frequencies. In other words, a relatively long LPC analysis window can be used for relatively low frequency bands to capture the longer periodicity of such signals. The LPC analysis parameters provide a flexible means for controlling the quantization noise shaping characteristics of the encoder.
[0146] Embodiments of a method for estimating the tonality of an audio signal will be described later in connection with FIG. 6.
[0147] In some embodiments, the weighting of the first and second SVs is based on the tonality estimate value T. T is a continuous variable ranging from 0 for a pure noise signal to 1 for a pure sine wave and a sparse harmonic signal component. Thus, the first and second SVs may be mixed as a linear combination using weights that sum to 1, and the weights depend on T. In other words, weighting the first and second SVs based on at least one characteristic of the audio signal includes calculating a value representing the weight, the value being in the range of 0 to 1, and the step of calculating the SV for the frequency band S08 includes multiplying one of the first and second SVs by the value, multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums to form the SV for the frequency band S08.
[0148] It should be understood that the function for calculating the SV S08 can be modified in various ways by scalar modification.
[0149] In one embodiment, the analysis component 704 is configured to use a third scalar when calculating the second SV S12.
[0150] As an example, the second SV may be calculated S12 by multiplying the difference by a third scalar different from the first scalar. It should be understood that different values may be assigned to the third scalar. The third scalar may be, for example, a value in the range of 0.05 to 0.2. The third scalar may be in the range of 0.1 to 0.15.
[0151] In one embodiment, the analysis component 704 is configured to use a fourth scalar when calculating the second SV S12.
[0152] As an example, the second SV may be calculated S12 by adding a fourth scalar to the difference multiplied by a third scalar, where the fourth scalar is different from the second scalar. It should be understood that different values may be assigned to the fourth scalar. The fourth scalar may be, for example, in the range of 3.5 to 4.5. The fourth scalar is typically set according to a threshold value in a quiet state.
[0153] Note that the second and fourth scalars may vary significantly depending on the setting for the threshold value in a quiet state. An important aspect of these terms is to allow a trade-off in the number of bits assigned to the tonal signal versus the noise-like signal. They are also useful for calibrating the model such that the noise assigned to exactly the level and shape of the masking threshold is barely perceptible to the average listener.
[0154] In one embodiment, the analysis component is configured to calculate S12 the second SV by multiplying the difference by 0.15 and adding 4 to the result.
[0155] Then, as an example, the overall SV may be calculated S08 as a weighted combination of the SV rules for, for example, a pure sine wave and a pure noise signal: SV(j)=max(0,T*(0.32*(Eb(j)-Q(j))+3)+(1-T)*(0.13*(Eb(j)-Q(j))+4))
[0156] FIG. 5 shows a model of the SV versus the perceived level for three different signal types. The tonal SV model 502 shows the model behavior for a signal with T = 1. The noise SV model 504 shows the model behavior for a signal with T = 0. The mixed tone and noise SV model 506 shows the model behavior for the case of T = 0.65.
[0157] Thus, the analysis component 700 can be configured to blend between the tone masking model and the noise masking model. In other words, for very tone-like signals, the encoder mainly uses a configuration suitable for tone-like signals. For very noise-like signals, the encoder mainly uses a configuration suitable for noise-like signals. For intermediate signals, the encoder uses a blend of those configurations, and the ratio of the tone-like configuration to the noise-like configuration depends on the in-band tonality within the band.
[0158] Returning to FIG. 7, then, a masking threshold may be calculated S16 using the sensitivity value and the energy value, which can then be used, in combination with the energy value, to determine S18 the bit allocation value for that frequency band.
[0159] Advantageously, the analysis component 700 calculates S16 the masking threshold by subtracting a variable offset (sensitivity value) from the signal energy or a value calculated based on the signal energy. The variable offset is based on, for example, the difference (sensation level) between the energy value and the auditory threshold at rest, as described above. In particular, as the sensation level increases, the variable offset increases, and vice versa. Such a way of calculating the masking threshold provides a better match to subjectively measured masking data and thus results in an improved allocation of bits. The improvement in the subjective quality of the decoded audio signal can be most pronounced for higher level signals. Prior art models using an offset independent of level generate a lower masking threshold than necessary for quieter signals, resulting in an over-allocation of bits and thus reducing the number of available bits for other bands and other frames containing louder signal components.
[0160] For comparison, prior art models typically simply determine the masking threshold by subtracting a fixed offset from the in-band signal energy. For example, in some cases, the same offset is used regardless of how close the band energy is to the auditory threshold. The analysis component 700, instead, determines the masking threshold by subtracting a variable offset from the signal energy.
[0161] The masking threshold can be calculated S16 in various ways. In one embodiment, calculating the masking threshold includes applying a spreading function to either the linear energy values for the frequency bands or the transformed energy values of the frequency bands. In other words, in one embodiment, the spreading function is applied to the energy values for the frequency bands. In another embodiment, the energy values are first transformed before the spreading function is applied. The transformation can include transforming the linear energy values into the loudness domain by raising the energy values to the power of about 0.25 to 0.3. Alternatively, the transformation can include raising the energy values to the power of 0.5 to 0.6, which has been found to provide better sound quality for some audio formats.
[0162] Thereby, an excitation value for that frequency band is determined. The excitation value is then combined with a sensitivity value to calculate the masking threshold. On the dB scale, the combination of the sensitivity value and the excitation value includes subtracting the sensitivity value from the excitation value. In the intensity domain, division is used instead.
[0163] In another embodiment, the spreading function is applied after combining the energy value and the sensitivity value to determine an intermediate threshold. In this embodiment, calculating the masking threshold includes combining the energy value and the sensitivity value to determine an intermediate threshold and applying the spreading function to the intermediate threshold to determine the masking threshold.
[0164] Optionally, for all of the above embodiments, the masking threshold is derived by threshold processing using the threshold at rest, for example, masking threshold = max(masking threshold, auditory threshold at rest).
[0165] In one embodiment, the spreading function for a certain frequency band depends on the sensation level, and the effect of the spreading function in a frequency band having a relatively high sensation level is greater compared to the effect of the spreading function in a frequency band having a relatively low sensation level. Typically, the spreading function is defined on an absolute SPL scale. Using an alternative method for defining the spreading function can provide a more generalized psychoacoustic model while incurring minimal additional computational complexity. At present, many encoders appear to apply the spreading function most suitable for quiet signals. This is a conservative design approach, but the degree of frequency-domain masking will be underestimated for louder signals. This may lead to allocating more bits than necessary in certain bands, and correspondingly, the number of bits available for other bands will be reduced, resulting in a potential quality degradation. Therefore, the analysis component 700 may be configured to define the spreading function for that frequency band depending on the difference between the energy value determined in decision S04 and the auditory threshold at rest determined in decision S06, which leads to an improvement in bit allocation.
[0166] In some embodiments, determining the bit allocation value S18 for a certain frequency band includes calculating the SMR for that frequency band, which is the energy value for that frequency band subtracted by the calculated masking threshold for that frequency band in calculation S16. In some embodiments, a further fixed offset is subtracted. Then, the determination S18 of the bit allocation value is based on the amount of SMR. In some embodiments, the bit allocation value is threshold processed with a defined maximum bit allocation value, for example, 12 bits.
[0167] In some embodiments, determining S18 the bit allocation values includes adjusting S20 a masking threshold to achieve a bit allocation that satisfies a target bit rate for the audio signal. Adjusting S20 the masking threshold may include adjusting the masking threshold by adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied. As described above, the conversion from the linear energy region to the loudness region includes raising each energy to the power of about 0.25 to 0.3.
[0168] Generally, the analysis component 700 allocates S18 more bits for a frequency band having a higher SV (when the SV is defined in dB as a subtractive adjustment to the excitation function) compared to when the frequency band has a lower SV.
[0169] In some embodiments, the analysis component 700 may be implemented in the encoder 800. Such an embodiment is shown in FIG. 8. In this embodiment, the encoder 800 includes a receiving component 802 configured to receive an audio signal 806. The encoder further includes an encoding component 804 configured to use the bit assignment value determined S18 by the analysis component 700 for encoding purposes. For example, the encoding component 804 is configured to quantize audio samples of the audio data in a frequency band in response to the bit assignment value and encode the quantized audio data in the frequency band into a bitstream 808. In some embodiments, the encoder 800 further includes a companding component (not shown) configured to reduce the dynamic range of the audio signal using a companding algorithm before quantizing the audio samples of the audio data in the frequency band. The companding function reduces the dynamic range of the input signal before transform coding. The companding function can be beneficial for the encoded quality of signals containing high-density transient mixed sounds such as rain and applause. In one example, the input signal companding and the embodiment of calculating S10 only the first SV and using S08 this as the SV may act synergistically to produce higher performance than each function separately. In this embodiment, the dynamic range of the audio signal is reduced using a companding algorithm before the step of encoding the audio signal using the relevant SV. The companding function can further reduce the number of bits to be encoded while still maintaining a high audio quality.
[0170] In some embodiments, the analysis component 700 is implemented in the decoder 900. This embodiment is shown in FIG. 9. In this embodiment, the decoder 900 includes a receiving component 902 configured to receive an audio signal 906 in the form of an encoded bitstream that includes encoded energy values for frequency bands of the audio signal. The decoder further includes a decoding component 904 configured to use the bit assignment values determined S18 by the analysis component 700 for decoding purposes. The decoding component 904 is configured to decode the encoded energy values from the encoded bitstream 906, and the analysis component 700 uses this decoded energy value when determining the energy value. The decoding component 904 is further configured to extract quantized audio samples of the audio data of that frequency band from the encoded bitstream 906 in response to the bit assignment value. The decoding component 904 is further configured to dequantize the quantized audio samples of the audio data of that frequency band and combine the dequantized audio samples of the audio data of each frequency band to generate a decoded audio signal 908.
[0171] Note that the analysis component 700 and the corresponding method can be used with any audio format.
[0172] The masking using the method of the present invention described herein provides a better match to subjectively measured masking data, thus resulting in an improvement in bit allocation. Embodiments that use the calculated S10 first SV as the SV for that frequency band provide the greatest improvement over the default encoder for speech signals. This is important because speech signals are a very important element of typical broadcast and movie content.
[0173] In some embodiments, the encoder 800 and / or decoder 900 implementing this embodiment (or alternatively, the embodiment where the second SV also calculates S12) are implemented in a real-time two-way communication device. Advantageously, this simpler embodiment can be used in such a device due to the lower complexity of such an encoding method. However, it should be noted that there are many applications and possible uses for the encoder 800 and / or decoder 900.
[0174] Thus, the encoder 800 accurately captures the observed masking behavior of the human auditory system. This leads to a higher codec performance than the default encoder for both constant bitrate and variable bitrate applications.
[0175] In some embodiments, the subjective improvement may be most pronounced for relatively high-level signals. This is because the default encoder that derives the masking threshold based on a level-independent offset (instead of using the SV) tends to over-allocate bits to low-level signal components.
[0176] FIG. 4 shows an overview of an exemplary method for calculating the masking threshold described above, where the tonality estimate is used to weight the first and second SVs. As shown in FIG. 4, an input audio frame is input to the MDCT filter bank. The transform length(s) input to the MDCT filter bank are also received by a tonality estimation unit (described further below in relation to FIG. 6). The tonality estimation unit outputs a tonality estimate T j (m) in the range of 0 to 1, and outputs one set for each MDCT transform. In this nomenclature, j is the band index and m is the MDCT block index.
[0177] The MDCT transform coefficients are used to determine the energy values for each frequency band. A spreading function is applied to the energy values of the various frequency bands to derive an excitation function. In the final steps of the exemplary method of FIG. 4, as described herein, the energy values and the quiet-time threshold are used to calculate the first and second SVs, which are then weighted by j (m) and applied to the excitation values to ultimately generate a masking threshold for each frequency band.
[0178] Here, an embodiment of the tone estimation method based on adaptive prediction will be described in connection with FIG. 6.
[0179] Input frame 602 provides input samples. Filter bank 604 is configured to receive input samples from input frame 602. Note that there are different filter banks 604 available. In one example, CMDCT is used with N = 128. In another example, CQMF may be used with N = 64. Filter bank 604 is configured to send complex frequency coefficients 606(X k (n), band k at time n) to LPC analysis component 608 and prediction improbability estimation component 605. The structure of FIG. 6 is repeated for each CMDCT / CQMF band. LPC analysis component 608 is configured to provide a set of prediction improbability values 609(μ k (n)) corresponding to frequency coefficients 606 in relation to prediction improbability estimation component 605.
[0180] Considering the fact that the time samples within one CMDCT block affect three adjacent CMDCT blocks, a 3-tap FIR filter is used to smooth the prediction improbability estimates in two-stage smoothing stage 620. This improves the smoothness of the tone estimation value (and thus the decoded audio as well). A similar approach is used for other filter banks, such as CQMF with N = 64.
[0181] The mapping component 610 is configured to receive a smoothed prediction improbability value, a transform length 612, and the energy of a frequency band (calculated by box 611 in FIG. 6). After the prediction improbability estimates are smoothed, they are combined across time to reflect the non-zero portions of the MDCT window to which they will later be applied. This can be important, especially for dynamically changing signals such as speech, in order to maximize the clarity of the decoded output signal. The mapping component 610 is further configured to send the mapped input data 613 (Z k (n)) as output data to the spreading and normalization component 614. The spreading and normalization component 614 is configured to apply a spreading function to the input data 613 and send a set of modified data 615 (U k (n)) to the tonal mapping component 616. The tonal mapping component 616 is configured to map the input data 615 (prediction improbability) to one or more sets of tonal estimates 618.
[0182] Moving on to a more detailed description of FIG. 6, according to some embodiments, a sine window is applied to a 50% overlapping block of input samples taken from one 4096-length frame 602, followed by the application of a 128-point CMDCT. The choice of filter bank 604 is not critical; for example, a complex QMF filter bank 604 already existing within a known encoder can also be used. For a set of complex frequency coefficients 606 X k (n) k = 1,…,N from block 604, a corresponding set of prediction improbability values 609 is generated. The prediction improbability for band k at time n is
Equation
[0183] At each CMDCT frequency bin k, L k consecutive coefficients X k (n - m), m = 1, …, L k are windowed, and analyzed to generate complex prediction coefficients of degree p k (p k < L k ). Then, prediction coefficients a ki , i = 1, …, p k are used to calculate the prediction impossibility value 609 corresponding to X k (n). Among the various LPC analysis windows evaluated, the nearly symmetric Hamming window was found to maximize the prediction gain. The degree of asymmetry varies as a function of the CMDCT bin. Then, the full-band prediction impossibility value is filtered by a set of two-stage smoothing filters 620 to avoid rapid changes over time. An exemplary two-stage filter consists of a normal exponential smoothing filter followed by a 3-tap FIR. The FIR filter receives the prediction impossibility value 609 μ k (n) and generates a partially smoothed output signal μ k '(n). Then, the FIR output signal is further processed by the exponential smoothing filter to generate μ k "(n).
[0184] In one embodiment, a high-speed attack low-speed decay IIR filter may be used for the exponential smoothing filter. These filters provide means for independently controlling the attack and decay times. The input is in the tonal region (1 - μ k '(n)), and the difference equation is given by:
Equation
[0185] In the absence of smoothing, the tone value tends to vary through successive conversion blocks, leading to fluctuations in the masking threshold estimate. This can lead to audible quantization noise modulation in the decoder output, especially at low to medium frequencies. An effective way to solve this problem is to design the attack / decay filter coefficients from the known time characteristics of the human auditory filter. This approach generally leads to the longest attack / decay time constants at low frequencies and the shortest at high frequencies.
[0186] In the next stage of Figure 6, the CMDCT bin energies and the smoothed unpredictability values from all CMDCT blocks are resampled and combined (mapped) into group 610 to match the time and frequency resolution of each MDCT transform within the current frame. The resampled unpredictability values 613 are weighted according to their relative perceptual importance and combined over time as needed. Exemplary perceptual weightings include the square of the L2 norm (energy) and loudness. Next, the unpredictability values are spread over frequency, for example, by applying a spreading function 614 from ISO / IEC 11172-3:1993(E) to match the spreading applied to the banded MDCT energies. In the final step, the resampled, spread, and normalized unpredictability values 615 are mapped to one or more sets of tone values 618 in the range from 0 to 1 (both ends inclusive) (one set for each MDCT transform).
[0187] Figure 10 shows results from examples of experimentally measured SVs for JND as a function of various tone + narrowband noise mixing ratios (SNRs). Center frequencies of 500 Hz 804, 1 kHz 802, and 4 kHz 806 were used with masker SNRs in the range from -10 dB to 40 dB. The masker was presented to the subject at a level of 80 dB SPL. The average curve 808 (thick line) represents the ta SV average across all three center frequencies.
[0188] In one embodiment, another tonal mapping function (different from the tonal mapping rules in ISO / IEC11172-3:1993) is calibrated based at least in part on the results of a perceptual masking experiment for a mixed tone + narrowband noise signal. The purpose of this embodiment is to determine the JND levels for a masker consisting of a tone + narrowband noise mixture and a maskee consisting of uncorrelated narrowband noise at the same frequency. The experiment is repeated at various masker - tone / noise mixing levels and various frequencies. The results of the experiment can be used to calibrate the tonal mapping as described below.
[0189] In one embodiment, first, each of the tone + narrowband noise stimuli is injected into a tone estimator to capture the associated prediction uncertainty values. From these results, a table is generated that associates each prediction uncertainty value with the required SMR. By combining this table with the tonal - to - SMR rules for tone + narrowband noise masking that match the SMR range, points on a curve that define the prediction - uncertainty - to - tonality mapping required for calibration of the model are derived. In a final step, a parametric function that approximates the derived calibration curve is derived. The masking experiment and calibration steps may be repeated at various frequencies and input signal levels.
[0190] FIG. 11 shows, by way of example, the results of an experiment for an embodiment. This presents the target calibration curve 904 and a parametric model / function 902 (dashed-dotted line) that approximates the derived calibration curve for a 128-point CMDCT using a sine window and 50% overlap. Examples of the prior art are included in the figure for comparison (dashed line 906). In this example, the LPC analysis is cubic with a window length of 6. In this embodiment, the parametric function T(μ) maps unpredictability to tonality as follows: T(μ)=min(max(a*μ 3 +b*μ 2 +c*μ+d,0),1)
[0191] The values for the four parameters (a, b, c, d) are derived to approximate the target calibration curve. In the example shown in FIG. 11, the model-parameter values for a, b, c, d are -13.0233, 15.9513, -8.1012, 2.1319.
[0192] The use of tonality estimation in a perceptual model is well known in the prior art (ISO / IEC 11172-3:1993(E), dashed line 906 in FIG. 11), but the prior art models operate in a level-independent manner. In ISO / IEC 11172-3:1993(E), the SMR model adapts based on an estimated tonality between about 6 and 30 dB for pure noise and tone signals, respectively. The level-dependent model described herein using a calibrated tonality mapping function has been implemented in simulations and has performance superior to that of the prior art models as determined by objective quality metrics and listening tests.
[0193] Further embodiments of the present disclosure will be apparent to those skilled in the art after considering the above description. Although this specification and the drawings disclose embodiments and examples, the present disclosure is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure as defined by the appended claims. Even if reference signs appear in the claims, they are not to be construed as limiting the scope.
[0194] Furthermore, variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the present disclosure, from a study of the drawings, the present disclosure, and the appended claims. In the claims, the term "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be advantageously used.
[0195] The systems and methods disclosed above can be implemented as software, firmware, hardware, or combinations thereof. In a hardware implementation, the division of tasks among the functional units described above does not necessarily correspond to a division into physical units. Conversely, one physical component may have multiple functions, or one task may be performed collaboratively by several physical components. Certain components or all components may be implemented as software executed by a digital signal processor or a microprocessor, or as hardware, or as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium that may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those skilled in the art, the term "computer storage medium" includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Further, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media as is well known to those skilled in the art.
[0196] Various aspects of the present disclosure may be understood from the following numbered example embodiments (EEE). 〔EEE1〕 A method for processing an audio signal, wherein the audio signal includes audio data in a plurality of frequency bands, and the method comprises: For each of the plurality of frequency bands: Determine an energy value for the audio data in that frequency band; Determine a quiet-time auditory threshold for that frequency band; Calculate a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold; Calculate a masking threshold for that frequency band using the sensitivity value and the energy value; Determining a bit allocation value for the frequency band using the energy value and the masking threshold, Method. 〔EEE2〕 Calculating the masking threshold comprises: The energy value for the frequency band; or The transformed energy value of the frequency band Applying a diffusion function to one of them to determine an excitation value for that frequency band; Combining the sensitivity value with the excitation value, The method according to EEE1. 〔EEE3〕 Calculating the masking threshold includes combining the energy value and the sensitivity value to determine an intermediate threshold, and applying a diffusion function to the intermediate threshold to determine the masking threshold, the method according to EEE1. 〔EEE4〕 The method according to any one of EEE1 to 3, further comprising quantizing audio samples of the audio data in the frequency band in response to the bit allocation value. 〔EEE5〕 The method according to EEE4, further comprising encoding the quantized audio data of the frequency band into a bit stream. 〔EEE6〕 The audio signal is an encoded bitstream including encoded energy values for the frequency band, and determining the energy value for the audio data of the frequency band includes decoding the encoded energy value from the encoded bitstream, the method according to any one of EEE1 to 3. 〔EEE7〕 The method according to EEE6, further comprising extracting quantized audio samples of the audio data of the frequency band from the encoded bitstream in response to the bit allocation value. 〔EEE8〕 The method according to EEE6, further comprising dequantizing the quantized audio samples of the audio data of the frequency band and combining the dequantized audio samples of the audio data of each frequency band to generate a decoded audio signal. 〔EEE9〕 The method according to any one of EEE1 to 8, wherein determining the bit allocation value includes adjusting the masking threshold to achieve a bit allocation that satisfies a target bit rate for the audio signal. 〔EEE10〕 The method according to EEE9, wherein adjusting the masking threshold includes adjusting the masking threshold by adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied. 〔EEE11〕 The method according to any one of EEE1 to 10, wherein the energy value, the quiet-time auditory threshold, and the masking threshold are defined in units of decibels (dB). 〔EEE12〕 The method according to any one of EEE1 to 11, further comprising determining the plurality of frequency bands of the audio signal according to an equivalent rectangular bandwidth (ERB) scale. 〔EEE13〕 The SV is defined in dB as a subtractive adjustment to the excitation value, and the step of determining the bit allocation value includes allocating more bits for a frequency band having a higher SV than for a frequency band having a lower SV, according to any one of EEE1 to 12 when citing EEE2 or EEE2. 〔EEE14〕 The step of calculating the SV for the frequency band includes calculating a first SV using a sensory level, which is the difference on a dB scale between the energy value and the auditory threshold at rest, according to any one of EEE1 to 13. 〔EEE15〕 The method according to EEE14, wherein the step of calculating the first SV includes multiplying the sensory level by a first scalar. 〔EEE16〕 The method according to EEE15, wherein the first scalar is frequency-dependent. 〔EEE17〕 The method according to EEE15, wherein the first scalar is constant across all frequency bands. 〔EEE18〕 The method according to any one of EEE15 to 17, wherein the step of calculating the first SV includes adding a second scalar to the sensory level multiplied by the first scalar. 〔EEE19〕 The method according to any one of EEE14 to 18, wherein the step of calculating the SV includes using the first SV as the SV for the frequency band. 〔EEE20〕 The method according to any one of EEE14 to 18, wherein the step of calculating the SV for the frequency band includes calculating a second SV using the sensory level and weighting the first and second SVs based on at least one characteristic of the audio signal. 〔EEE21〕 The method according to EEE20, wherein the step of calculating a second SV for the frequency band includes multiplying a third scalar, which is different from the first scalar, by the perceived level. 〔EEE22〕 The method according to EEE21, wherein the step of calculating a second SV includes adding a fourth scalar to the perceived level multiplied by the third scalar, and the fourth scalar is different from the second scalar. 〔EEE23〕 The method according to any one of EEE20 to EEE22, wherein the step of weighting the first and second SVs based on at least one characteristic of the audio signal includes calculating a value representing the weight, the value being in the range between 0 and 1, and the step of calculating the SV for the frequency band includes multiplying the value by one of the first and second SVs and multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums to form the SV for the frequency band. 〔EEE24〕 The method according to any one of EEE20 to EEE23, wherein the at least one characteristic defines an estimated tonality of the frequency band of the audio signal. 〔EEE25〕 The method according to any one of EEE20 to EEE23, wherein the at least one characteristic defines an estimated level of noise in the frequency band of the audio signal. 〔EEE26〕 The method according to EEE24, wherein the estimated tonality is calculated using an adaptive prediction of frequency coefficients calculated from the frequency band of the audio signal. 〔EEE27〕 The method according to EEE26, wherein linear predictive coding LPC is applied adaptively to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated. 〔EEE28〕 The method according to EEE27, wherein the LPC analysis window length is varied as a function of the frequency band. 〔EEE29〕 The method according to EEE28, wherein a relatively longer LPC analysis window is used for a relatively lower frequency band. 〔EEE30〕 The method according to any one of EEE27 to 29, wherein the prediction order of the LPC is varied as a function of the frequency band. 〔EEE31〕 The method according to any one of EEE1 to 30, wherein the frequency range of the audio signal is between 200 and 7000 Hz. 〔EEE32〕 The method according to any one of EEE1 to 31, including using a predefined table that defines the auditory threshold for at least some frequencies to determine the auditory threshold at rest for that frequency band. 〔EEE33〕 The method according to EEE4 or any other EEE when citing EEE4, wherein the dynamic range of the audio signal is reduced using a companding algorithm before quantizing the audio samples of the audio data in the frequency band. 〔EEE34〕 The method according to any one of EEE15 to 33 when citing EEE14 or EEE14, further including defining a diffusion function for that frequency band depending on the sensory level such that the effect of the diffusion function in a frequency band having a relatively higher sensory level is greater compared to the effect of the diffusion function in a frequency band having a relatively lower sensory level. 〔EEE35〕 A receiving component configured to receive an audio signal, wherein the audio signal includes audio data in a plurality of frequency bands, and an analysis component; An apparatus having an analysis component configured to determine the plurality of frequency bands of the audio signal, wherein the analysis component for each frequency band of the plurality of frequency bands: Determine an energy value for the audio data in the frequency band; Determine a quiet-time auditory threshold for the frequency band; Calculate a sensitivity value (SV) for the frequency band using the energy value and the quiet-time auditory threshold; Calculate a masking threshold for the frequency band using the sensitivity value and the energy value; Further configured to determine a bit allocation value for the frequency band using the energy value and the masking threshold, Device. 〔EEE36〕 The analysis component is: An energy value for the frequency band; or The transformed energy value of the frequency band Determine an excitation value for the frequency band by applying a spreading function to one of them; By combining the sensitivity value with the excitation value, The device according to EEE35, configured to calculate the masking threshold. 〔EEE37〕 The analysis component is configured to calculate the masking threshold by combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold, the device according to EEE35. 〔EEE38〕 An encoder, the device according to any one of EEE35 to 37, further having an encoding component configured to quantize audio samples of the audio data in the frequency band in response to the bit allocation value. 〔EEE39〕 The encoding component is further configured to encode the quantized audio data in the frequency band into a bit stream, the device according to EEE38. 〔EEE40〕 An apparatus according to any one of EEE35 to 37, which is a decoder, wherein the audio signal is an encoded bitstream including an encoded energy value for the frequency band, and the apparatus further has a decoding component configured to decode the encoded energy value from the encoded bitstream, and the analysis component uses the decoded energy value when determining the energy value. 〔EEE41〕 The apparatus according to EEE40, wherein the decoding component is further configured to extract quantized audio samples of the audio data in the frequency band from the encoded bitstream in response to the bit allocation value. 〔EEE42〕 The apparatus according to EEE41, wherein the decoding component further dequantizes the quantized audio samples of the audio data in the frequency band and combines the dequantized audio samples of the audio data in each frequency band to generate a decoded audio signal. 〔EEE43〕 The apparatus according to any one of EEE35 to 42, wherein the analysis component is configured to adjust the masking threshold so as to achieve a bit allocation that satisfies a target bit rate for the audio signal when determining the bit allocation value. 〔EEE44〕 The apparatus according to EEE43, wherein the analysis component is configured to adjust the masking threshold by adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied when adjusting the masking threshold. 〔EEE45〕 The analysis component is configured to define the energy value, the auditory threshold at rest, and the masking threshold in units of decibels (dB), and the device according to any one of EEE35 to EEE44. 〔EEE46〕 The analysis component is configured to determine the plurality of frequency bands of the audio signal according to an equivalent rectangular bandwidth (ERB) scale, and the device according to any one of EEE35 to EEE45. 〔EEE47〕 The SV is defined in dB as a subtractive adjustment to the excitation value, and the analysis component is configured to determine a bit allocation value by allocating more bits for a frequency band having a higher SV than for a frequency band having a lower SV, and the device according to any one of EEE36 or EEE37 to EEE46. 〔EEE48〕 The analysis component is configured to calculate the SV for the frequency band by calculating a first SV using a sensory level, where the sensory level is the difference on a dB scale between the energy value and the auditory threshold at rest, and the device according to any one of EEE35 to EEE47. 〔EEE49〕 The analysis component is configured to calculate the first SV by multiplying the sensory level by a first scalar, and the device according to EEE48. 〔EEE50〕 The first scalar is frequency-dependent, and the device according to EEE49. 〔EEE51〕 The first scalar is constant across all frequency bands, and the device according to EEE49. 〔EEE52〕 The analysis component is configured to calculate the first SV by adding a second scalar to the sensory level multiplied by the first scalar, and the device according to any one of EEE49 to EEE51. 〔EEE53〕 The apparatus according to any one of EEE48 to 52, wherein the analysis component is configured to calculate the SV by using the first SV as the SV for its frequency band. 〔EEE54〕 The apparatus according to any one of EEE48 to 52, wherein the analysis component is configured to calculate a second SV by using the perceived level, and to calculate the SV for its frequency band by weighting the first and second SVs based on at least one characteristic of the audio signal. 〔EEE55〕 The apparatus according to EEE54, wherein the analysis component is configured to calculate the second SV for its frequency band by multiplying a third scalar different from the first scalar by the perceived level. 〔EEE56〕 The apparatus according to EEE55, wherein the analysis component is configured to calculate the second SV by adding a fourth scalar to the perceived level multiplied by the third scalar, and the fourth scalar is different from the second scalar. 〔EEE57〕 The apparatus according to any one of EEE54 to 55, wherein the analysis component performs the steps of calculating a value representing the weight for weighting the first and second SVs based on at least one characteristic of the audio signal, the value being in the range between 0 and 1, multiplying one of the first and second SVs by the value, multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums to form the SV for its frequency band. 〔EEE58〕 The apparatus according to any one of EEE54 to 57, wherein the at least one characteristic defines the estimated tonality of the frequency band of the audio signal. 〔EEE59〕 The at least one characteristic defines an estimated level of noise in the frequency band of the audio signal, the apparatus according to any one of EEE54 to 57. [EEE60] The analysis component is configured to calculate the estimated tonality using an adaptive prediction of frequency coefficients calculated from the frequency band of the audio signal, the apparatus according to EEE58. [EEE61] The analysis component is configured to adaptively apply LPC to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated, the apparatus according to EEE60. [EEE62] The LPC analysis window length is varied as a function of the frequency band, the apparatus according to EEE61. [EEE63] A relatively longer LPC analysis window is used for a relatively lower frequency band, the apparatus according to EEE62. [EEE64] The prediction order of the LPC is varied as a function of the frequency band, the apparatus according to any one of EEE62 to 63. [EEE65] The frequency range of the audio signal is between 200 and 7000 Hz, the apparatus according to any one of EEE35 to 64. [EEE66] Further comprising a memory, the memory storing a table defining the hearing threshold at rest for at least some frequencies, the analysis component being configured to determine the hearing threshold at rest for the frequency band by using the predefined table, the apparatus according to any one of EEE35 to 65. [EEE67] An apparatus according to any other EEE when citing EEE38 or EEE38, further comprising a companding component configured to reduce the dynamic range of the audio signal using a companding algorithm before quantizing audio samples of the audio data in the frequency band. 〔EEE68〕 The apparatus according to any one of EEE48 or EEE49 to 67, wherein the analysis component is configured to define a diffusion function for the frequency band depending on the perceived level such that the effect of the diffusion function in a frequency band having a relatively higher perceived level is greater compared to the effect of the diffusion function in a frequency band having a relatively lower perceived level. 〔EEE69〕 An apparatus according to any one of EEE35 to 67, implemented in a real-time two-way communication device. 〔EEE70〕 A method for estimating the tonality of an input signal, comprising: applying a filter bank to achieve a set of frequency coefficients; and calculating an estimated tonality using an adaptive prediction of the frequency coefficients. Method. 〔EEE71〕 The method according to EEE70, wherein the step of calculating the estimated tonality includes applying adaptive linear prediction to the frequency coefficients based on the frequency band of the audio signal from which the frequency coefficients are calculated. 〔EEE72〕 The method according to EEE70, wherein the LPC analysis window length is varied as a function of the frequency band. 〔EEE73〕 The method according to EEE72, wherein a relatively longer LPC analysis window is used for a relatively lower frequency band. 〔EEE74〕 The method according to any one of EEE72 to 73, wherein the prediction order of the LPC is varied as a function of the frequency band. 〔EEE75〕 The filter bank includes one of a 128-band complex MDCT or DFT filter bank and a 64-band complex QMF filter bank, according to the method of any one of EEE70 to 74. [EEE76] The LPC analysis window is an asymmetric Hamming window, according to the method of any one of EEE71 to 73. [EEE77] The method according to any one of EEE70 to 76, including the step of weighting the predictability index from the adaptive prediction according to the relative perceptual importance of each predictability index. [EEE78] The step of weighting the predictability index included in each time-frequency tile includes one of weighting based on the energy or loudness of the input signal, according to the method described in EEE77. [EEE79] The method according to any one of EEE70 to 78, further including the step of combining the predictability indices from the adaptive prediction of the frequency coefficients so as to match the time and frequency resolution of the filter bank. [EEE80] The sensitivity value and the excitation value are defined in decibels (dB), and the combining step includes subtracting the sensitivity value from the excitation value, or the sensitivity value and the excitation value are defined on an intensity scale, and the combining step includes calculating the quotient of the excitation value and the sensitivity value, according to the method of any one of EEE4 to 34 when citing EEE2 or EEE2. [EEE81] The energy value and the sensitivity value are defined in decibels (dB), and the combining step includes subtracting the sensitivity value from the energy value, or the energy value and the sensitivity value are defined on an intensity scale, and the combining step includes calculating the quotient of the energy value and the sensitivity value, according to the method of any one of EEE4 to 34 when citing EEE3 or EEE3. [EEE82] Calculating the sensitivity value includes calculating a ratio or difference between the energy value of the frequency band and the auditory threshold value at rest in the frequency band, according to any one of EEE1 to 34 or EEE80 to 81. 〔EEE83〕 A computer program product including a computer-readable storage medium having instructions adapted to execute the method according to any one of EEE1 to 34 or EEE80 to 82 when executed by a device having processing capabilities. 〔EEE84〕 A computer program product including a computer-readable storage medium having instructions adapted to execute the method according to any one of EEE70 to 70 when executed by a device having processing capabilities. 〔EEE85〕 The sensitivity value and the excitation value are defined in decibels (dB), and the combining step includes subtracting the sensitivity value from the excitation value, or the sensitivity value and the excitation value are defined on an intensity scale, and the combining step includes calculating the quotient of the excitation value and the sensitivity value, according to any one of EEE36, or EEE38 to 69 when citing EEEE36. 〔EEE86〕 The energy value and the sensitivity value are defined in decibels (dB), and the combining step includes subtracting the sensitivity value from the energy value, or the energy value and the sensitivity value are defined on an intensity scale, and the combining step includes calculating the quotient of the energy value and the sensitivity value, according to the device described in EEE37, or any one of EEE38 to 69 when citing EEEE37.
Claims
**Claim 1** A method for processing an audio signal, executed by one or more processors, wherein the audio signal includes audio data in a plurality of frequency bands, the method comprising: For each of the plurality of frequency bands: Determining an energy value for the audio data of that frequency band; Determining a quiet-time auditory threshold for that frequency band; Calculating a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold, wherein calculating the sensitivity value includes calculating a ratio or difference between the energy value of that frequency band and the quiet-time auditory threshold for that frequency band; Calculating a masking threshold for that frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold comprises: applying a spreading function to one of the energy value for that frequency band or a transformed energy value of that frequency band to determine an excitation value for that frequency band; and combining the sensitivity value with the excitation value; Determining a bit allocation value for the frequency band using the energy value and the masking threshold. A method. **Claim 2** A method for processing an audio signal, executed by one or more processors, wherein the audio signal includes audio data in a plurality of frequency bands, the method comprising: For each of the plurality of frequency bands: Determining an energy value for the audio data of that frequency band; Determining a quiet-time auditory threshold for that frequency band; Calculating a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold, wherein calculating the sensitivity value includes calculating a ratio or difference between the energy value of that frequency band and the quiet-time auditory threshold for that frequency band; A step of calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes determining an intermediate threshold by combining the energy value and the sensitivity value, and determining the masking threshold by applying a diffusion function to the intermediate threshold; A step of determining a bit allocation value for the frequency band using the energy value and the masking threshold; Method.
3. The method according to claim 1 or 2, wherein when the calculated masking threshold is greater than the calculated masking threshold, the calculated masking threshold is replaced with the calculated masking threshold.
4. Determining the bit allocation value includes adjusting the masking threshold to achieve a bit allocation that satisfies a target bit rate for the audio signal, Adjusting the masking threshold includes adding a constant offset to the masking threshold in the loudness region until the target bit rate for the audio signal is satisfied, thereby adjusting the masking threshold. The method according to any one of claims 1 to 3.
5. The method according to claim 3 when depending on claim 1 or claim 4 when depending on claim 1, wherein the SV is defined in dB as a subtractive adjustment to the excitation value, and the step of determining the bit allocation value includes allocating more bits for a frequency band having a higher SV than for a frequency band having a lower SV.
6. The method according to any one of claims 1 to 5, wherein the step of calculating the SV for the frequency band includes calculating a first SV using a sensation level, which is the difference on a dB scale between the energy value and the calculated masking threshold.
7. The method according to claim 6, wherein the step of calculating the first SV includes multiplying the sensation level by a first scalar, and / or the step of calculating the SV includes using the first SV as the SV for the frequency band.
8. The method according to claim 6 or 7, wherein the step of calculating the SV for the frequency band includes calculating a second SV using the perceived level and weighting the first and second SVs based on at least one characteristic of the audio signal.
9. The method according to claim 8, wherein the at least one characteristic defines an estimated level of tonality in the frequency band of the audio signal.
10. The method according to claim 9, wherein the estimated level of tonality is calculated using an adaptive prediction of frequency coefficients calculated from the frequency band of the audio signal.
11. The method according to claim 10, wherein linear predictive coding LPC is applied adaptively to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated.
12. The LPC analysis window length is varied as a function of the frequency band, The method according to claim 11, wherein the prediction order of the LPC is varied as a function of the frequency band.
13. The method according to any one of claims 6 to 12, wherein the spreading function for the frequency band depends on the perceived level such that the effect of the spreading function in a frequency band having a relatively higher perceived level is greater compared to the effect of the spreading function in a frequency band having a relatively lower perceived level.
14. The method according to any one of claims 1 to 13, further comprising quantizing audio samples of the audio data of the frequency band in response to the bit allocation value; and encoding the quantized audio data of the frequency band into a bitstream.
15. The method according to claim 14, wherein the dynamic range of the audio signal is reduced using a companding algorithm before quantizing the audio samples of the audio data of the frequency band.
16. The method according to any one of claims 1 to 13, wherein the audio signal is an encoded bitstream including an encoded energy value for the frequency band, and determining the energy value for the audio data for the frequency band includes decoding the encoded energy value from the encoded bitstream.
17. In response to the bit allocation value, extracting quantized audio samples of the audio data in the frequency band from the encoded bitstream; dequantizing the quantized audio samples of the audio data in the frequency band, and combining the dequantized audio samples of the audio data in each frequency band to generate a decoded audio signal, the method according to claim 16, further comprising.
18. A receiving component configured to receive an audio signal, the audio signal including audio data in a plurality of frequency bands, an analysis component; An apparatus having an analysis component configured to determine a plurality of frequency bands of the audio signal, The analysis component, for each frequency band of the plurality of frequency bands: Determining an energy value for the audio data in that frequency band; Determining a quiet-time auditory threshold for that frequency band; Calculating a sensitivity value (SV) for that frequency band using the energy value and the quiet-time auditory threshold, calculating the sensitivity value including calculating a ratio or difference between the energy value in that frequency band and the quiet-time auditory threshold for that frequency band, step; Calculating a masking threshold for that frequency band using the sensitivity value and the energy value, calculating the masking threshold comprising: applying a spreading function to one of the energy value for that frequency band or the transformed energy value for that frequency band to determine an excitation value for that frequency band; combining the sensitivity value with the excitation value, step; Determining a bit allocation value for the frequency band using the energy value and the masking threshold, further configured to perform. Apparatus.
19. A receiving component configured to receive an audio signal, the audio signal including audio data in a plurality of frequency bands, an analysis component; An apparatus having an analysis component configured to determine a plurality of frequency bands of the audio signal, The analysis component, for each frequency band of the plurality of frequency bands: Determining an energy value for the audio data in the frequency band; Determining a quiet-time auditory threshold for the frequency band; Calculating a sensitivity value (SV) for the frequency band using the energy value and the quiet-time auditory threshold, wherein calculating the sensitivity value includes calculating a ratio or difference between the energy value of the frequency band and the quiet-time auditory threshold of the frequency band; Calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes determining an intermediate threshold by combining the energy value and the sensitivity value, and applying a spreading function to the intermediate threshold to determine the masking threshold; Determining a bit allocation value for the frequency band using the energy value and the masking threshold; and Apparatus.
20. A computer program for causing one or more processors to execute the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Method and device for allocating dynamic bit for audio coding
JP2000004163A