Psychoacoustic model for audio processing
By calculating masking thresholds using sensitivity values and energy values, the method addresses inefficiencies in audio encoding, enhancing audio quality and bit allocation accuracy for diverse audio content.
Patent Information
- Application Number
- JP2025080876
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-05
AI Technical Summary
Existing audio encoding technologies fail to accurately estimate masking thresholds based on human hearing properties, leading to inefficient bit allocation and degraded audio quality, particularly for audio content with dynamic level changes.
A method for calculating masking thresholds using sensitivity values (SV) and energy values in frequency bands, incorporating the human auditory system's response to adjust excitation functions, thereby improving bit allocation and audio quality.
The method provides more accurate bit allocation, resulting in higher quality audio encoding across varying audio content, including speech and music, with reduced complexity and improved consistency.
Smart Images

Figure 2025114804000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 943,903, filed December 5, 2019, and European Patent Application No. 19213742.0, filed December 5, 2019, both of which are incorporated herein by reference in their entireties.
[0002] Technical Field The present disclosure relates to the field of audio processing, and in particular to a method for processing an audio signal using a masking model based on hearing thresholds in quiet for frequency ranges of the audio signal and measured energy values of the audio signal for the corresponding frequency ranges, and further to an apparatus capable of performing the audio processing method. [Background technology]
[0003] The human brain cannot register all audio signals at all different frequencies. Therefore, when coding audio, it is beneficial to remove signals at frequencies and levels that are imperceptible to the human auditory system. This is typically done by removing insignificant components from the audio signal. Within the context of perceptual audio coders, there are two main ways that encoders can increase compression efficiency: removing signal redundancy and insignificance. Redundant (predictable) signal components are typically removed in the encoder and restored in the decoder. Insignificant signal components are typically removed in the audio encoder by quantization and are not restored by the audio decoder.
[0004] Typically, encoders use a psychoacoustic model, sometimes called a perceptual model, to estimate a masking threshold for the audio spectrum. The masking threshold provides an estimate of the minimum just-noticeable distortion (JND) that can be tolerated in each of several frequency bands of the audio spectrum. Frequency bands are typically nonuniform in width, according to the critical bands of the human auditory system. In a typical encoder, the masking threshold is input to a rate control loop that selects scale factors (and quantization noise levels) for each of several scale factor bands. The performance of a typical encoder depends on how close the masking threshold estimate is to the true JND noise level: a masking threshold estimate that exceeds the JND noise level results in fewer bits being allocated than necessary to avoid audible distortion, while a masking threshold estimate that falls below the JND noise level results in more bits being allocated than necessary, potentially at the expense of neighboring frequency bands.
[0005] Typically, the encoder determines the masking threshold by: 1) Calculate the signal energy of the audio signal on the critical band frequency scale. 2) The frequency response of the signal after processing by the basilar membrane (also known as the excitation function) is estimated by convolving the critical band energy with a set of spreading functions. 3) For each critical band, adjust the excitation function for that band downward by the amount estimated to achieve the JND noise level.
[0006] Furthermore, models for determining masking thresholds often include heuristic rules that are developed experimentally but are not directly based on known properties of human hearing.
[0007] Thus, there is room for improvement in the art of calculating masking thresholds based on known properties of human hearing in order to improve bit allocation for frequency bands of an audio signal. Summary of the Invention [Problem to be solved by the invention]
[0008] In view of the above, it is an object of the present disclosure to overcome or mitigate at least some of the above-mentioned problems. In particular, it is an object of the present disclosure to provide a masking model based on the energy value of an audio signal in a frequency band and the hearing threshold in quiet for that frequency band. Furthermore, it is an object of the present disclosure to provide a masking model that reduces the complexity of audio coding and improves the quality of encoded audio based on the above. Further and / or alternative objects of the present disclosure will be apparent to the reader of this disclosure. [Means for solving the problem]
[0009] According to a first aspect, there is provided a method of processing an audio signal, the audio signal comprising audio data in a plurality of frequency bands, the method comprising: For each frequency band of the plurality of frequency bands: Determine the energy value for the audio data in that frequency band; Determine the hearing threshold in quiet for that frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet; calculating a masking threshold for the frequency band using the sensitivity value and the energy value; determining a bit allocation value for the frequency band using the energy value and the masking threshold.
[0010] With respect to the term "energy value", in the context of this specification, it should be understood that various approaches for calculating these energies may be used. The calculation may be based, for example, on a banded modified discrete cosine transform (MDCT), a discrete Fourier transform (DFT), or a complex MDCT (CMDCT). It should be noted that several energy values may be calculated for a frequency band and then combined in an appropriate manner to form a single energy value for that frequency band. In this specification, "energy value" may refer to energy expressed either linearly or in dB scale.
[0011] With respect to the term "frequency band," in the context of this specification, it should be understood that a frequency band is an interval in the frequency domain, bounded by lower and upper frequency limits and having a frequency range. It should be noted that the frequency bands of an audio signal to be encoded do not necessarily have the same width / range. For example, a relatively low frequency band may have a width of 100-200 Hz, and a relatively high frequency band may have a width of 3000-3500 Hz. Typically, the width of a frequency band increases with increasing frequency, and the frequency band between the relatively low and high frequency bands may typically have a width somewhere in the range of 100-3000 Hz.
[0012] The term "sensitivity value" (SV) in the context of this specification should be understood as an approximation of the adjustment required in a given critical band to achieve a JND distortion for a human listener with normal hearing. To take into account the effects of masking across the critical band, the SV for each band may depend not only on the signal characteristics within that band but also on signals in neighboring bands. The SV for each band is typically applied as an offset or adjustment to the excitation function, and then a threshold in quiet is applied to derive the final masked threshold. Any noise below the masked threshold is inaudible.
[0013] The SV for a particular frequency band can be calculated, for example, using the ratio or difference between the energy value of that frequency band and the hearing threshold in quiet for that frequency band, or any other metric that compares the energy value with the hearing threshold in quiet.
[0014] In typical prior art encoders, downward adjustments made to the excitation function in critical frequency bands are typically invariant to signal level, except for the application of a quiet threshold at the end. As a result, estimated masking thresholds may not correlate perfectly with the masking behavior of the human auditory system.
[0015] Thus, the representation of tuning for JNDs is typically level-independent. Such models are typically based on masking data for relatively loud or relatively quiet signals. This approach can limit codec performance, for example, by underestimating the true JND threshold for low-level signal components and resulting in an over-allocation of bits to frames containing relatively quiet signal passages. This problem arises not only for variable bit-rate encoders, but also for encoders operating in constant bit-rate mode with bit basing. Audio content characterized by highly dynamic level changes, such as speech, is adversely affected.
[0016] In the present disclosure, by calculating the masking threshold based on both SV and energy values, the masking threshold can more accurately capture the observed masking behavior of the human auditory system, thereby delivering a higher quality audio signal.
[0017] Furthermore, when encoding an audio signal using a model that more faithfully captures the observed masking behavior of the human auditory system, the method can more accurately estimate the number of bits required to meet a predetermined quality target to provide a constant quality audio signal, thereby reducing bit over- or under-allocation. In embodiments where a constant bit rate is desired, the method can provide an improved quality audio signal due to an improved bit allocation strategy.
[0018] The method may further provide a better match to subjectively measured masking data. Using the described audio encoding model, a single model appropriate for all recorded sound levels or audio content may be achieved. Advantageously, the model may facilitate encoding of audio signals of consistent quality, regardless of the characteristics of the audio signal to be encoded. Some examples of audio signal characteristics are pitch, loudness, or duration, although it should be noted that there are many other characteristics of audio signals.
[0019] General Description of the Embodiments According to some embodiments, calculating the masking threshold comprises applying a spreading function to one of the energy values for the frequency bands; or the transformed energy values of the frequency bands to determine excitation values for the frequency bands; Combining the sensitivity value with the excitation value.
[0020] The excitation function can be thought of as the energy distribution along the basilar membrane of the inner ear, and the excitation value is then the value calculated from that function for a particular frequency band.
[0021] To mimic the processing of sound in the basilar membrane of the ear and smooth the predictability index across frequency, a spreading function is applied to the energy values or the transformed energy values. For example, the spreading function may be applied to the energy values transformed into the loudness domain (i.e., to raise the energy values to a power of approximately 0.25 to 0.3). In other embodiments, the spreading function may be applied to the energy values raised to a power of 0.5 to 0.6. Spreading functions from ISO / IEC 11172-3:1993(E) may be used.
[0022] If the sensitivity value and the excitation value are defined in decibels (dB), the combining step may include calculating the masking threshold by subtracting the sensitivity value from the excitation value. On an intensity scale, the masking threshold is calculated as the quotient of the excitation value and the sensitivity value.
[0023] Optionally, the masked threshold is derived by thresholding with the threshold in quiet, for example masked threshold=max(masked threshold, hearing threshold in quiet).
[0024] According to some embodiments, calculating the masking threshold comprises combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold.
[0025] For example, the masking threshold may be determined as max(median threshold, hearing threshold in quiet).
[0026] According to some embodiments, the method further comprises quantizing audio samples of the audio data for the frequency bands in response to the bit allocation value. Advantageously, the encoder can encode the audio with a constant quality or with improved audio quality at a constant bit rate. The encoder can further encode the quantized audio data for the frequency bands into a bitstream.
[0027] The methods described herein may also be used on the decoder side. According to some embodiments, the audio signal is an encoded bitstream including encoded energy values for frequency bands, and determining the energy values for the audio data of the frequency bands includes decoding the encoded energy values from the encoded bitstream. On the decoder side, the determined bit allocation values may be used to extract quantized audio samples of the audio data of the frequency bands from the encoded bitstream. Advantageously, the bit allocation values for each frequency band of each audio frame need not be included in the bitstream but can instead be determined on the decoder side. The bitrate of the encoded bitstream may be reduced in this way.
[0028] According to some embodiments, the method further includes dequantizing the quantized audio samples of the audio data for the frequency bands and combining the dequantized audio samples of the audio data for each frequency band to generate a decoded audio signal.
[0029] According to some embodiments, determining the bit allocation value includes adjusting the masking threshold to achieve a bit allocation that meets the target bit rate for the audio signal. In this embodiment, if the number of bits required by the nominal masking threshold is more (or less) than the number of bits available to meet the bit rate requirements, the masking threshold can be adjusted to allocate more or fewer bits to use as many bits as possible without exceeding the target bit rate. For example, adjusting the masking threshold may include adjusting the masking threshold by adding a constant offset to the masking threshold in the loudness domain until the target bit rate for the audio signal is met.
[0030] Different measurements can be used when determining and defining energy and hearing threshold. According to some embodiments, energy values, hearing threshold in quiet and masked threshold are defined in decibels dB. This provides simplicity to the model as decibels are a common measure of loudness / energy.
[0031] According to some embodiments, the method comprises determining the plurality of frequency bands of the audio signal according to an Equivalent Rectangular Bandwidth (ERB) scale. The ERB scale provides an approximation of the bandwidth of the human auditory system using the convenient simplification of modeling the auditory filter as a rectangular bandpass filter. Advantageously, using ERB can be beneficial when encoding an audio signal according to the human auditory system.
[0032] According to some embodiments, the SV is defined in dB as a subtractive adjustment to the excitation function, and the step of determining the bit allocation value comprises allocating more bits to frequency bands with higher SV than to frequency bands with lower SV. Advantageously, a consistent audio quality of the encoded audio signal can be achieved. The SV controls the displacement of the excitation function that gives the masking threshold after application of a quiet threshold. Positive sensitivity values lower the masking threshold, while negative sensitivity values raise the masking threshold. Consequently, an increasing sensitivity value corresponds to a lower masking threshold and thus to more allocated bits. Thus, the sensitivity value for a frequency band can be considered to correspond to the sensitivity of the human auditory system to noise (coding artifacts) in that frequency band of the audio signal.
[0033] According to some embodiments, the step of calculating the SV for the frequency band comprises calculating a first SV using a sensation level, the sensation level being the difference in dB scale between the energy value and the hearing threshold in quiet.
[0034] The term "difference" is to be understood in the context of this document as the energy value (expressed in dB) minus the hearing threshold in quiet (expressed in dB).
[0035] The term sensation level as used herein is defined as the level of a sound relative to the threshold in quiet for an average listener. This term was introduced by [1]. [Non-Patent Document 1] C.J. Moore, "An Introduction to the Psychology of Hearing", 5th edition, p. 403, Academic Press (2003)
[0036] It should be noted that determining the exact SV can be accomplished in a variety of ways.
[0037] According to some embodiments, the step of calculating the first SV comprises multiplying the sensory level by a first scalar. Advantageously, greater accuracy of the SV can be achieved with less complexity relative to the human auditory system. By multiplying the function by a first scalar, the difference and the SV can be more easily mapped to each other to better correspond to the human auditory system.
[0038] The first scalar may be frequency dependent or constant across all frequency bands.
[0039] According to some embodiments, calculating the first SV includes adding a second scalar to the sensation level multiplied by the first scalar, where the second scalar may be frequency dependent or constant across all frequency bands.
[0040] Advantageously, higher accuracy of the SV with respect to the human auditory system can be achieved with lower complexity: by adding a second scalar to the function, the mapping between the difference and the SV can be easily modified to better correspond to the human auditory system.
[0041] According to some embodiments, the step of calculating the SV includes using the first SV as the SV for the frequency band.
[0042] According to some embodiments, the step of calculating the SV for the frequency band includes calculating a second SV using the sensation level and weighting the first and second SV based on at least one characteristic of the audio signal.
[0043] By way of example, such characteristics of an audio signal can be bandwidth, tone-to-noise ratio, nominal level, or power level in decibels (dB). However, it should be noted that an audio signal has many characteristics that can be used in this method. By weighting the first and second SVs, a masking threshold close to the JND can be obtained. Advantageously, a constant audio quality can be achieved, regardless of the audio content of the audio signal.
[0044] In this embodiment, when using the above model in which two SVs are calculated and weighted, a noisy, loud audio signal can result in a high-quality encoded audio signal. A less noisy, soft audio signal can also result in the same level of quality for the encoded audio signal when using the same model. Conversely, in prior art dual-mode encoders, for audio signals that cannot be reliably classified, such as dialogue mixed with applause, the encoder must select either the applause mode or the default mode, and the mode selected by the encoder may not be optimal. Alternatively, the encoder may choose to encode different segments of the signal in different modes, which may result in audible switching artifacts that degrade the quality of the encoded signal. In a single-mode model in which first and second SVs are calculated and weighted, it is not necessary to determine which portions of the audio signal should be encoded in different modes. Furthermore, a single-mode model is suitable for diverse audio signals that do not depend on the audio content (e.g., speech, music, etc.) of the audio signal. By providing a single model, it is not necessary to classify audio samples to determine which mode should be applied for encoding. Furthermore, the problem of mode selection for signals at the boundary between modes is alleviated, and audio quality degradation due to the encoder making suboptimal mode selection is avoided.
[0045] Note that in some embodiments, further SVs can be defined and included in calculating the final SV (by weighting more than two SVs). These different SVs may be weighted based on the transient characteristics of the audio signal.
[0046] According to some embodiments, calculating the second SV for the frequency band includes multiplying the sensation level by a third scalar different from the first scalar, which may be frequency dependent or constant across all frequency bands.
[0047] Multiplying the function by a third scalar can improve the accuracy of the SV relative to the auditory system for audio signals with different characteristics, such as audio signals with a high degree of noise-like or tone-like characteristics. The third scalar can allow the SV to be mapped to enable a generalized model for audio coding. The SV is defined as a function of the sensory level, as understood above and further described below. By mapping the SV according to different characteristics of the audio signal, the relationship between the hearing threshold and the energy value can remain unchanged, while the slope of the SV and the sensory level can change according to different characteristics of the audio signal. Advantageously, bits can be allocated to provide encoded audio signals with high quality.
[0048] According to some embodiments, calculating the second SV includes adding a fourth scalar to the sensation level multiplied by the third scalar, the fourth scalar being different from the second scalar. Note that the fourth scalar may be assigned various values to improve the accuracy of the SV with respect to the human auditory system. The fourth scalar may be frequency dependent or constant across all frequency bands.
[0049] According to some embodiments, weighting the first and second SVs based on at least one characteristic of the audio signal comprises calculating a value representing a weight, the value ranging from 0 to 1, and calculating the SV for the frequency band comprises multiplying one of the first and second SVs by the value and multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums together to form the SV for the frequency band.
[0050] In other words, the first SV and the second SV are mixed as a linear combination with weights that sum to one, the weights being dependent on said at least one characteristic of the audio signal.
[0051] By calculating an overall SV by weighting the first and second SVs, the mapping between the perceptual level and the SV can be adapted to reflect the masking characteristics of different audio signal types. Advantageously, a low-complexity and flexible model for determining bit allocation values for frequency bands can be achieved, which provides a high-quality encoded audio signal.
[0052] According to some embodiments, said at least one characteristic defines an estimated tonality of that frequency band of the audio signal.
[0053] It should be noted that there are many different ways to estimate the tonality of an audio signal. Tonality describes the relationship between tonal characteristics (e.g., notes, chords, keys, pitches, etc.) in an audio signal. Advantageously, using the estimated tonality of that frequency band of the audio signal as a characteristic for weighting the first and second SVs may improve the accuracy of the SVs relative to the human auditory system. Furthermore, using tonality may improve subjective audio quality.
[0054] According to some embodiments, the at least one characteristic defines an estimated level of noise in that frequency band of the audio signal, which noise may advantageously be masked relative to the human auditory system in order to achieve a high quality encoded audio signal.
[0055] According to some embodiments, the estimated tonality is calculated using adaptive prediction of frequency coefficients calculated from that frequency band of the audio signal. In some embodiments, the same set of frequency coefficients is used for both the masking threshold calculation and the tonality estimation. In other embodiments, the tonality estimation is performed using a separate complex-valued filter bank. Note that any set of frequency coefficients is possible, depending on the desired accuracy of the estimated tonality and the available computational resources. For example, using only real MDCT coefficients is computationally cheaper than using CMDCT coefficients, but is less accurate. Obtaining an accurate tonality estimate can further improve the subjective audio quality.
[0056] According to some embodiments, linear predictive coding (LPC) is adaptively applied to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated. LPC may be used rather than fixed prediction to achieve a more accurate tonality estimate of the audio signal.
[0057] It should be noted that the LPC analysis window may have different lengths. By varying the analysis window length, a desired variable time-frequency framework can be flexibly realized. According to some embodiments, the LPC analysis window length is varied as a function of the frequency band. In some embodiments, a relatively longer LPC analysis window is used for a relatively lower frequency band.
[0058] According to some embodiments, the LPC prediction order is varied as a function of frequency band. For example, the LPC prediction order may be selected to maximize the discrimination between a pure noise input and a signal with a tonal component (such as a harpsichord or speech).
[0059] It should be noted that the frequency range of the audio signal may be coded into different ranges.
[0060] According to some embodiments, the frequency range of the audio signal is between 200 and 7000 Hz.
[0061] According to some embodiments, the step of determining the hearing threshold in quiet for a frequency band comprises using a predefined table defining said hearing threshold for at least some frequencies. The predefined table can be pre-stored in the encoder performing the method, thereby allowing the predefined table to be updated without affecting decoder compatibility. Advantageously, this can reduce the complexity of providing a high-quality encoded audio signal.
[0062] According to some embodiments, before quantizing audio samples of the frequency band audio data, the dynamic range of the audio signal is reduced using a companding algorithm. By companding the audio signal in the encoder and applying a complementary expansion in the decoder, the encoding method can provide a higher quality decoded audio signal. Companding the audio signal allows for fewer bits to be coded while maintaining high audio quality.
[0063] According to some embodiments, the method includes a step of defining a spreading function for a frequency band depending on the sensation level, such that the effect of the spreading function in that frequency band is greater than the effect of the spreading function in a frequency band having a relatively high sensation level compared to the effect of the spreading function in a frequency band having a relatively low sensation level.
[0064] According to a second aspect, at least one of the above objects is achieved by an apparatus comprising: an analysis component configured to receive an audio signal, the audio signal including audio data in a plurality of frequency bands; an analysis component configured to determine a plurality of frequency bands of the audio signal; The analysis component, for each frequency band of the plurality of frequency bands: Determine the energy value for the audio data in that frequency band; Determine the hearing threshold in quiet for that frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet; calculating a masking threshold for the frequency band using the sensitivity value and the energy value; The device is configured to determine a bit allocation value for the frequency band using the energy value and the masking threshold.
[0065] According to some embodiments, the analysis component: the energy values of those frequency bands; or The transformed energy values for those frequency bands determining an excitation value for the frequency band by applying a spreading function to one of the By combining the sensitivity value with the excitation value, The masking threshold is calculated.
[0066] According to some embodiments, the analysis component is configured to calculate the masking threshold by combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold.
[0067] According to some embodiments, the apparatus is an encoder, and the apparatus further comprises an encoding component configured to quantize audio samples of the audio data for frequency bands in response to the bit allocation values.
[0068] According to some embodiments, the encoding component is further configured to encode the quantized audio data of the frequency bands into a bitstream.
[0069] According to some embodiments, the apparatus is a decoder, and the audio signal is an encoded bitstream comprising encoded energy values for frequency bands, and the apparatus further comprises a decoding component configured to decode the encoded energy values from the encoded bitstream, and the analysis component uses the decoded energy values when determining the energy values.
[0070] According to some embodiments, the decoding component is further configured to extract quantized audio samples of frequency band audio data from the encoded bitstream in response to the bit allocation value.
[0071] According to some embodiments, the decoding component is further configured to dequantize the quantized audio samples of the audio data for the frequency bands and combine the dequantized audio samples of the audio data for each frequency band to generate a decoded audio signal.
[0072] According to some embodiments, the analysis component, when determining the bit allocation value, is configured to adjust the masking threshold to achieve a bit allocation that meets a target bit rate for the audio signal.
[0073] According to some embodiments, the analysis component is configured to, when adjusting the masking threshold, adjust the masking threshold by adding a constant offset to the masking threshold in the loudness domain until a target bitrate for the audio signal is met.
[0074] According to some embodiments, the analysis component is configured to define the energy values, hearing thresholds in quiet and masked thresholds in decibels dB.
[0075] According to some embodiments, the analysis component is configured to determine the plurality of frequency bands of the audio signal according to an Equivalent Rectangular Bandwidth (ERB) scale.
[0076] According to some embodiments, the SV is defined in dB as a subtractive adjustment to the excitation function, and the analysis component is configured to determine a bit allocation value by allocating more bits for frequency bands with higher SV compared to frequency bands with lower SV.
[0077] According to some embodiments, the analysis component is configured to calculate an SV for the frequency band by calculating a first SV using a sensation level, the sensation level being the difference between the energy value and the hearing threshold in quiet.
[0078] According to some embodiments, the analysis component is configured to calculate the first SV by multiplying the sensation level by a first scalar.
[0079] According to some embodiments, the first scalar is frequency dependent.
[0080] According to some embodiments, the first scalar is constant across all frequency bands.
[0081] According to some embodiments, the analysis component is configured to calculate the first SV by adding a second scalar to the sensation level multiplied by the first scalar.
[0082] According to some embodiments, the analysis component is configured to calculate the SV by using the first SV as the SV for that frequency band.
[0083] According to some embodiments, the analysis component is configured to further calculate a second SV using the sensation level and to calculate an SV for the frequency band by weighting the first and second SV based on at least one characteristic of the audio signal.
[0084] According to some embodiments, the analysis component is configured to calculate the second SV for the frequency band by multiplying the sensation level by a third scalar different from the first scalar.
[0085] According to some embodiments, the analysis component is configured to calculate the second SV by adding a fourth scalar to the sensation level multiplied by the third scalar, wherein the fourth scalar is different from the second scalar.
[0086] According to some embodiments, the analysis component is configured to weight the first and second SVs based on at least one characteristic of the audio signal by calculating a value representing the weight, the value ranging from 0 to 1, multiplying one of the first and second SVs by the value and multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums together to form the SV for the frequency band.
[0087] According to some embodiments, said at least one characteristic defines an estimated tonality of that frequency band of the audio signal.
[0088] According to some embodiments, the at least one characteristic defines an estimated level of noise in that frequency band of the audio signal.
[0089] According to some embodiments, the analysis component is configured to calculate the estimated tonality using adaptive prediction of frequency coefficients calculated from that frequency band of the audio signal.
[0090] According to some embodiments, the analysis component is configured to adaptively apply LPC to the MDCT coefficients based on a frequency band of the audio signal, the frequency band from which the MDCT coefficients are calculated.
[0091] According to some embodiments, the LPC analysis window length is varied as a function of frequency band.
[0092] According to some embodiments, a longer LPC analysis window is used for the lower frequency bands.
[0093] According to some embodiments, the prediction order of the LPC is varied as a function of frequency band.
[0094] According to some embodiments, the frequency range of the audio signal is between 200 and 7000 Hz.
[0095] According to some embodiments, the device further comprises a memory storing a table defining hearing thresholds in quiet for at least some frequencies, and the analysis component is configured to determine hearing thresholds in quiet for frequency bands by using the predefined table.
[0096] According to some embodiments, the apparatus further comprises a companding component configured to reduce the dynamic range of the audio signal using a companding algorithm before quantizing the audio samples of the audio data for the frequency bands.
[0097] According to some embodiments, the analysis component is configured to define a spreading function for a frequency band depending on the sensation level, such that the effect of the spreading function in a frequency band having a relatively high sensation level is greater compared to the effect of the spreading function in a frequency band having a relatively low sensation level.
[0098] According to some embodiments, the device is implemented as a real-time two-way communication device.
[0099] The second aspect may generally have the same advantages as the first aspect.
[0100] According to a third aspect, there is provided a method for estimating tonality of an input signal, comprising: applying a filter bank to achieve a set of frequency coefficients; and calculating an estimated tonality using adaptive prediction of the frequency coefficients.
[0101] According to some embodiments, the step of calculating the estimated tonality comprises applying adaptive linear prediction (LPC) to the frequency coefficients based on frequency bands of the audio signal from which the frequency coefficients are calculated.
[0102] According to some embodiments, the LPC analysis window length is varied as a function of frequency band.
[0103] According to some embodiments, a relatively long LPC analysis window is used for the relatively low frequency band.
[0104] According to some embodiments, the prediction order of the LPC is varied as a function of frequency band.
[0105] According to some embodiments, the filter bank includes one of a 128-band complex MDCT or DFT filter bank and a 64-band complex quadrature mirror filter CQMF filter bank.
[0106] According to some embodiments, the LPC analysis window is an asymmetric Hamming window.
[0107] According to some embodiments, the method comprises: Weighting the predictability measures from the adaptive prediction according to the relative perceptual importance of each predictability measure.
[0108] According to some embodiments, weighting the predictability measures contained within each time-frequency tile includes one of weighting based on the energy or loudness of the input signal.
[0109] According to some embodiments, the method comprises: The method further includes combining predictability measures from the adaptive prediction of frequency coefficients to match the time and frequency resolution of the filter bank.
[0110] Furthermore, it should be noted that unless expressly stated otherwise, the present disclosure relates to all possible combinations of features. [Brief explanation of the drawings]
[0111] The above, as well as additional objects, features, and advantages of the present disclosure, will be better understood through the following illustrative and non-limiting detailed description of embodiments of the present disclosure, with reference to the accompanying drawings, in which like reference numerals are used to refer to similar elements, in which: [Figure 1] Masking data for audio signals are shown. [Figure 2] Masking data for audio signals are shown. [Figure 3] 1 shows experimental data of signal-to-mask ratio SMR in relation to sensation level for a variety of tones at various frequencies, and a linear model of the data. [Figure 4] 1 shows an overview of a method for calculating a masking threshold according to some embodiments. [Figure 5] 10 shows SV in relation to sensation level for pure tones and pure noises according to some embodiments. [Figure 6] FIG. 1 shows a block diagram for estimating tonality for frequency bands of an input frame according to some embodiments. [Figure 7] 1 illustrates an analysis component for determining bit allocation values for frequency bands of an input audio signal according to some embodiments. [Figure 8] 8 shows an encoder that implements the analysis component of FIG. 7. [Figure 9] 8 shows a decoder that implements the analysis component of FIG. 7. [Figure 10] Results from an example of measured SV for JND as a function of tone / noise mixture level (SNR) are shown. [Figure 11] 7 shows a plot of estimated tonality and unpredictability estimates as defined in FIG. 6 in connection with the prior art. DETAILED DESCRIPTION OF THE INVENTION
[0112] The present disclosure will now be described more fully with reference to the accompanying drawings, in which embodiments of the disclosure are shown. The systems and devices disclosed herein will be described in operation.
[0113] In the following, a known audio format is used as a context for illustrating the present disclosure. However, it should be noted that the scope of the present disclosure is not limited to this known format, and the various embodiments described herein may be used with any suitable audio format.
[0114] For the exemplary format, there are currently two commonly used modes for encoding audio. Choosing which mode is most appropriate for an audio signal can be a complex decision, and if a mode that is not appropriate for the audio signal is selected, the quality of the encoded audio signal may be degraded. Two typical modes are default and applause. The current modes are distinct; in both modes, the encoder estimates the masking threshold from the signal energy estimate and the SV, which is invariant to the signal level, except for applying a quiet threshold at the end. The default mode also applies legacy functionality inherited from MPEG Layer III encoders, for which there is no sufficient perceptual justification. Furthermore, the masking threshold is input to a rate control loop that selects the scale factor (and quantization level) for each of several scale factor bands. Performance therefore depends on how close the masking threshold estimate is to the true JND noise level.
[0115] In most prior art models, the representation of the required SMR for the JND is level-independent before the application of the quiet threshold. Such models are typically based on masking data for relatively loud or relatively quiet signals, both of which are non-adaptive. This approach can, in some cases, limit codec performance by underestimating the true JND threshold for low-level signal components and resulting in excessive allocation of bits to frames containing relatively quiet signal passages. This problem arises not only for variable bitrate encoders, but also for encoders operating in constant bitrate mode with bit basing. Audio content characterized by highly dynamic level changes (e.g., speech) will be adversely affected.
[0116] A general problem with prior art models is that they produce lower masking thresholds than necessary, which also leads to an over-allocation of bits in frequency bands, thus reducing the number of bits available for other bands and thereby degrading the quality of the encoded audio signal.
[0117] The present disclosure aims to avoid some of the above problems by providing a single model that estimates a more accurate SMR and therefore performs as well or better for most audio content than prior art dual-mode or single models.
[0118] Subjective listening tests using mono content show that the new encoder outperforms current encoders for speech content. Additionally, the new encoder is significantly more effective in variable bitrate applications, where the encoder allocates only the number of bits necessary to meet a predefined quality target, providing constant audio quality.
[0119] In one experiment, to quantify the benefits of level-dependent masking, we conducted a first subjective listening test using three encoders and a diverse set of audio test items. Of the three encoders, one operated in default mode, one operated in applause mode, and one operated using level-dependent masking. The encoder using level-dependent masking provided an average subjective quality improvement of 3 and 14 points over the default and applause mode encoders, respectively. More importantly, level-dependent masking improved two speech items by an average of 8 points over the default encoder.
[0120] Figure 7 shows, by way of example, an analysis component 700. As further described below in connection with Figures 8 and 9, the analysis component may be implemented in the encoder 800 or the decoder 900. In other embodiments, the analysis component is implemented in a separate device, e.g., connected to the encoder or decoder.
[0121] The analysis component 700 includes circuitry configured to perform a method for processing an audio signal to determine bit allocation values for frequency bands of the audio signal. The circuitry may include one or more processors.
[0122] The analysis component 700 is configured to perform a variety of actions, examples of which are listed below.
[0123] The analysis component 700 is configured to determine (S02) multiple frequency bands of the input audio signal. Each of the multiple frequency bands includes a frequency range. It should be noted that each of the multiple frequency bands of the audio signal to be encoded does not necessarily have to have the same width / range. In one example, a first relatively low frequency band may have a range of 100-200 Hz, and another relatively high frequency band may have a range of 3000-3500 Hz. In one embodiment, the frequency range of the audio signal may be 200-7000 Hz. It should also be noted that an audio signal may have many different frequency ranges that may extend to frequencies above 7000 Hz and / or below 200 Hz. As will be appreciated, there are various ways to determine the frequency bands for an audio signal. In one embodiment, the analysis component 700 is configured to determine (S02) the frequency bands according to an equivalent rectangular bandwidth (ERB) scale. The ERB scale provides an approximation of the bandwidth of filters in the human auditory system. Furthermore, using the ERB scale provides the simplification of modeling the filter as a rectangular bandpass filter.
[0124] The analysis component is further configured to determine S18 a bit allocation value for each frequency band using the following analysis of the audio data for each frequency band.
[0125] The analysis component 700 determines S04 energy values for the audio data in frequency bands, which may be, for example, banded MDCT energies.
[0126] Further, analysis component 700 determines S06 hearing thresholds in quiet for the frequency bands. In some embodiments, analysis component 700 includes or is connected to a memory component. The memory component stores a table defining hearing thresholds in quiet for at least some frequencies. It should be noted that such a memory component can store different information. In other words, determining S06 hearing thresholds in quiet for the frequency bands can include using a predefined table defining hearing thresholds for at least some frequencies. In some embodiments, the predefined table defining hearing thresholds can be replaced, allowing for improvements to the encoder without affecting decoder compatibility.
[0127] Using the energy value and the hearing threshold in quiet, a sensitivity value (SV) can be calculated S08. It should be understood that the SV can be calculated S08 in various ways using the energy value and the hearing threshold in quiet. For example, the SV can be calculated S08 using the ratio or difference between the energy value and the hearing threshold in quiet, or any other metric that compares the energy value and the hearing threshold in quiet. A sensitivity value should be understood as a quantity defined, for example, in dB.
[0128] In one embodiment, the first SV is calculated S10 using the difference between the energy value and a hearing threshold in quiet (also referred to in this disclosure as the "sensation level"). Optionally, the first SV may be calculated S10 by multiplying the sensation level by a first scalar. In some embodiments, the first SV may be calculated S10 by adding a second scalar to the difference multiplied by the first scalar. Thus, in this embodiment, the first SV for a frequency band is calculated as alpha * (band energy - hthresh) + beta, where alpha is the first scalar, beta is the second scalar, band energy is the energy value of the audio signal for that frequency band, and hthresh is the threshold in quiet for that frequency band. In some embodiments, the second scalar is not included in the SV calculation to reduce complexity.
[0129] The degree to which the first SV varies with the difference between the energy value and the threshold in quiet for different frequency bands can be determined by examining a variety of measured masking data. Note that the following measurements and figures, described in connection with Figures 1-3, are provided using the difference between the energy value in each frequency band and the hearing threshold in quiet, i.e., the sensation level, as an example. However, those skilled in the art will understand that other data can be obtained experimentally if other ways of calculating SV are used, for example, using the ratio between the energy and the hearing threshold for each frequency band.
[0130] Figure 1 shows, as an example, measured masking data for a 200 Hz tone at different sound pressure levels (SPL). Masking thresholds 104, 106, 108, 110, and 112 are presented relative to the hearing threshold in quiet 102 (bold line). Masking threshold 1 104 relates to a 200 Hz tone at 60 dB SPL. Masking threshold 2 106 relates to a 200 Hz tone at 80 dB SPL. Masking threshold 3 108 relates to a 200 Hz tone at 90 dB SPL. Masking threshold 4 110 relates to a 200 Hz tone at 100 dB SPL. Masking threshold 5 112 relates to a 200 Hz tone at 105 dB SPL. As can be seen in Figure 1, the difference between the masked threshold at 200 Hz for various tone masker levels (i.e., 60, 80, 90, 100, and 105 dB; not specifically shown in Figure 1, but easily seen by tracing the vertical axis from each sound level to the 200 Hz marking on the horizontal axis) increases with increasing tone masker intensity. For a 60 dB tone masker, the difference is approximately 18 dB (tone masker 60 dB, masked threshold 104 is 42 dB), and for a 105 dB tone masker, the difference is approximately 32 dB (tone masker 105 dB, masked threshold 112 is 73 dB).
[0131] Figure 2 shows a similar pattern for the 500 Hz tone masker, where masking thresholds 204, 206, 208, and 210 for masking data at different sound pressure levels (SPL) (same levels in Figure 2 except that the 105 dB tone masker is not shown) are presented in relation to the hearing threshold in quiet 102 (bold line).
[0132] In one example, measured masking data (i.e., as illustrated in Figures 1 and 2) can be used to derive a sensitivity value, SV, versus sensation level when a tonal or sinusoidal signal is considered. It will be appreciated that there are other parameters than those illustrated in Figures 1-2 that can be used to derive the required SV. In this example, we considered masking for coding artifacts within the auditory critical band at a sinusoidal masker frequency.
[0133] Figure 3 shows, as an example, a model of SMR as a combination of tone-masking narrowband noise for different tone perception levels and frequencies. SMR is presented as a function of perception level using a linear model 302. Linear model 302 is a combination of measured SMR curves 1 through 6, designated 304, 306, 310, 312, 314, and 316 in Figure 3. Curve 1 304 shows the measured SMR value for a signal with a frequency of 200 Hz. Curve 2 306 shows the measured SMR value for a signal with a frequency of 500 Hz. Curve 3 310 shows the measured SMR value for another signal with a frequency of 500 Hz. Curve 4 312 shows the measured SMR value for a signal with a frequency of 1000 Hz. Curve 5 314 shows the measured SMR value for a signal with a frequency of 2000 Hz. Curve 6 316 shows the measured SMR value for a signal with a frequency of 5000 Hz.
[0134] Figure 3 shows that for this example, the slope of SMR vs. sensation level, 0.35 dB * (masker level relative to threshold in quiet at that frequency) + 3 dB, may be a reasonable approximation for the frequency range 200-4000 Hz. Note, however, that the linear model 302 in Figure 3 may be a reasonable approximation for other frequency ranges. In this example, the decibel offset for the required SMR varies by up to 10 dB at mid-levels but appears to converge at high and low levels.
[0135] In some embodiments, the quiet threshold may be modified by setting the threshold for all bands below 4 kHz to a global minimum threshold. The quiet threshold should be set to the minimum value within each band when encoding. For example, in a transform codec with adaptive block switching, the lowest frequency band of the shortest transform block may be 750 Hz wide. As can be seen from Figures 1-2, the hearing threshold level (dB) during quiet decreases rapidly from 20 to 750 Hz. The threshold for this entire band can then be set to the actual quiet threshold at 750 Hz. The same steps are applied to all other bands in the shortest block. These values are then interpolated to obtain quiet thresholds for all other transform block lengths. This approach ensures that the quiet threshold is at a consistent level for all block lengths and avoids undesirable quantization noise modulation artifacts when the codec switches transform lengths. An alternative, simpler approach is to set the threshold for all bands below 4 kHz to a global minimum threshold. Using this adjusted quiet threshold will result in other values for the first and / or second scalars, as will be understood by those skilled in the art.
[0136] Note that the quiet thresholds in Figures 1-2 are conservatively placed 20 dB below what would be expected under the typical assumption of a peak playback level of 105 dB SPL. In one embodiment, the thresholds are set based on a peak playback level of 115 dB. This provides a degree of robustness, especially for variable bit rate applications, when playing back decoded audio at levels different from the expected level.
[0137] The model in Figure 3 is derived by averaging the results of a tone-masking narrowband noise experiment for various frequencies. Signal components with higher perceptual levels experience higher SV. In one example, for every 3 dB increase in band energy above the hearing threshold in quiet, the SV increases by 1 dB. The level-dependent SV model SV(j) in Figure 3 for frequency band j is expressed as: SV(j)=max(0,0.35*(Eb(j)-Q(j))+3) where Eb(j) and Q(j) (expressed in dB in this example) are the banded MDCT energy and the quiet-time threshold, respectively.
[0138] As will be appreciated, changing the configuration of the analysis component will obviously change the scalars presented in the above equations. These scalars may be modified to adjust the SV calculation to better suit certain audio signals. The first scalar may be, for example, in the range of 0.2 to 0.5. The second scalar may be, for example, in the range of 2.5 to 3.5.
[0139] In the model of Figure 3, the first scalar and the second scalar are constant across all frequency bands, however, in other embodiments, the first and / or second scalars are frequency dependent.
[0140] In one embodiment, as shown in FIG. 3, a first SV is calculated S10 and used as the SV for that frequency band.
[0141] In one embodiment, the linear model 302 in FIG. 3 is extended to more accurately estimate the SV for input signals with high levels of noise. Some examples of input signals with high noise levels could be applause, or rain, or sibilant speech. However, as will be appreciated, there are many different signals with high levels of noise. For example, when using the same methodology as for tone-masking noise, the SV can be calculated to be more accurate for some signals.
[0142] In a second embodiment, we propose that at high noise levels, the same linear relationship exists between SV and sensation level, but with a different slope. The slope of the best-fit line is roughly half that of the case with tone-masking noise. This correspondence has been verified using experiments similar to those shown in Figures 1-3, but for noise maskers instead of tone maskers. Thus, a generalized model can be realized by adapting the SV rule depending on the input signal characteristics.
[0143] Thus, in some embodiments, an optional second SV is calculated S12 and combined with the first SV, optionally using a fixed or adaptive weighted combination, to define a final SV. In these embodiments, the analysis component 700 is further configured to use the difference between the determined S04 energy and the determined S06 hearing threshold in quiet (sensation level) to calculate S12 the second SV when calculating S08 the SV for the frequency band, and to weight the first and second SVs based on at least one determined S14 characteristic of the input audio signal. As will be appreciated, any suitable characteristic of the audio signal can be used in calculating S08 the SV. In some embodiments, the at least one characteristic is an estimated tonality of the signal. Alternatively, in some embodiments, the at least one characteristic is an estimated noise level for the signal.
[0144] In one embodiment, the estimated tonality is calculated using adaptive prediction of frequency coefficients calculated from a frequency band of the audio signal. In the following, an embodiment for estimating the tonality of an audio signal is described.
[0145] As will be appreciated, any set of frequency coefficients can be used. As an example, one prior art method is based on a second-order fixed prediction of the magnitude and phase of the DFT over time (Non-Patent Document 2). [Non-patent document 2] According to ISO / IEC 11172-3:1993(E), "Information technology - Coding of moving pictures and associated audio for digital storage media at up to about 1.5 Mbit / s - Part 3: Audio," overlapped DFTs of lengths 512 and 128 (i.e., number of complex DFT coefficients) are computed in parallel to allow for different time / frequency resolution tradeoffs for different frequencies. The analysis component 700 can generalize prior art methods to use adaptive linear prediction of the complex MDCT (CMDCT) coefficients. In some embodiments, linear predictive coding (LPC) can be adaptively applied to the MDCT coefficients based on the frequency band of the audio signal over which the MDCT coefficients are computed. Adaptive linear prediction allows for rapidly evolving mid-range harmonics in voiced speech and music to produce a higher tonality estimate than fixed prediction. Additionally, the desired variable time / frequency framework can be flexibly realized without the need for a parallel CMDCT filter bank by varying the LPC analysis window length and / or prediction order as a function of frequency. In other words, the LPC analysis window length can be varied as a function of frequency band. Furthermore, the LPC prediction order can also be varied as a function of frequency band. Optimal LPC analysis parameters may be selected offline for each frequency band by maximizing the difference between the average prediction gain for the challenging signal and independent and identically distributed (IID) Gaussian noise. Examples of challenging signals include speech or a harpsichord. However, it should be understood that many different signals can be classified as challenging. The longest LPC analysis window is typically used at low frequencies, with progressively shorter windows used at higher frequencies. In other words, relatively long LPC analysis windows may be used for relatively low frequency bands to capture the longer periodicity of such signals. The LPC analysis parameters provide a flexible means for controlling the quantization noise shaping characteristics of the encoder.
[0146] An embodiment of a method for estimating the tonality of an audio signal is described below in connection with FIG.
[0147] In some embodiments, the weighting of the first and second SVs is based on a tonality estimate T, where T is a continuous variable ranging from 0 for a pure noise signal to 1 for a pure sinusoidal and sparse harmonic signal component. Thus, the first and second SVs may be combined as a linear combination with weights that sum to 1, the weights being dependent on T. In other words, weighting the first and second SVs based on at least one characteristic of the audio signal includes calculating a value representing the weight, the value ranging from 0 to 1, and calculating S08 the SV for the frequency band includes multiplying one of the first and second SVs by the value and multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums together to form S08 the SV for the frequency band.
[0148] It should be understood that the function that calculates the SV S08 can be modified in various ways by modifying the scalar.
[0149] In one embodiment, the analysis component 704 is configured to use a third scalar in calculating S12 the second SV.
[0150] For example, the second SV may be calculated S12 by multiplying the difference by a third scalar different from the first scalar. It should be understood that the third scalar may be assigned a different value. The third scalar may be, for example, a value in the range of 0.05 to 0.2. The third scalar may be in the range of 0.1 to 0.15.
[0151] In one embodiment, the analysis component 704 is configured to use a fourth scalar in calculating S12 the second SV.
[0152] For example, a second SV may be calculated S12 by adding a fourth scalar to the difference multiplied by a third scalar, where the fourth scalar is different from the second scalar. It should be understood that the fourth scalar may be assigned a different value. The fourth scalar may be, for example, in the range of 3.5 to 4.5. The fourth scalar is typically set according to a quiet threshold.
[0153] Note that the second and fourth scalars can vary significantly depending on the setting for threshold in quiet. One important aspect of these terms is that they allow for a trade-off in the number of bits allocated to tonal signals versus noise-like signals. They are also useful for calibrating the model so that noise allocated to the correct masking threshold level and shape is barely perceptible to the average listener.
[0154] In one embodiment, the analysis component is configured to calculate S12 a second SV by multiplying the difference by 0.15 and adding 4 to the result.
[0155] An overall SV may then be calculated S08, for example as a weighted combination of SV rules for a pure sine wave and a pure noise signal: SV(j)=max(0,T*(0.32*(Eb(j)-Q(j))+3)+(1-T)*(0.13*(Eb(j)-Q(j))+4))
[0156] Figure 5 shows models of SV versus sensation level for three different signal types. Tonal SV model 502 shows model behavior for signals with T = 1. Noise SV model 504 shows model behavior for signals with T = 0. Mixed tonal and noise SV model 506 shows model behavior for T = 0.65.
[0157] Thus, analysis component 700 can be configured to blend between tone and noise masking models. In other words, for very tone-like signals, the encoder uses primarily the configuration appropriate for tone-like signals. For very noise-like signals, the encoder uses primarily the configuration appropriate for noise-like signals. For intermediate signals, the encoder uses a blend of the configurations, with the ratio of tone-like and noise-like configurations depending on the in-band tonality.
[0158] Returning to Figure 7, the sensitivity value and the energy value may then be used to calculate a masking threshold S16, which can then be used in combination with the energy value to determine a bit allocation value for that frequency band S18.
[0159] Advantageously, the analysis component 700 calculates S16 the masking threshold by subtracting a variable offset (sensitivity value) from the signal energy or a value calculated based on the signal energy. The variable offset, as described above, is based on, for example, the difference between the energy value and the hearing threshold in quiet (sensation level). In particular, as the sensation level increases, the variable offset increases, and vice versa. This manner of calculating the masking threshold provides a better match to subjectively measured masking data and thus results in improved bit allocation. The improvement in the subjective quality of the decoded audio signal may be most noticeable for signals with higher levels. Prior art models using a level-independent offset generate lower-than-necessary masking thresholds for quieter signals, resulting in over-allocation of bits and consequently reducing the number of available bits for other bands and frames containing louder signal components.
[0160] For comparison, prior art models typically simply determine the masking threshold by subtracting a fixed offset from the in-band signal energy. For example, in some cases, the same offset is used regardless of how close the in-band energy is to the hearing threshold. Analysis component 700 instead determines the masking threshold by subtracting a variable offset from the signal energy.
[0161] The masking threshold may be calculated S16 in various ways. In one embodiment, calculating the masking threshold includes applying a spreading function to either the linear energy values for the frequency bands or the transformed energy values for the frequency bands. In other words, in one embodiment, the spreading function is applied to the energy values for the frequency bands. In another embodiment, the energy values are first transformed before the spreading function is applied. The transformation may include converting the linear energy values to the loudness domain by raising the energy values to a power of approximately 0.25 to 0.3. The transformation may alternatively include raising the energy values to a power of 0.5 to 0.6, which has been found to provide better sound quality for some audio formats.
[0162] This determines an excitation value for that frequency band. The excitation value is then combined with the sensitivity value to calculate the masking threshold. In the dB scale, combining the sensitivity and excitation values involves subtracting the sensitivity value from the excitation value. In the intensity domain, division is used instead.
[0163] In another embodiment, the spreading function is applied after combining the energy value and the sensitivity value to determine an intermediate threshold value, in this embodiment, calculating the masking threshold comprises combining the energy value and the sensitivity value to determine an intermediate threshold value, and applying the spreading function to the intermediate threshold value to determine the masking threshold value.
[0164] Optionally, for all of the above embodiments, the masking threshold is derived by thresholding with the threshold in quiet, for example, masking threshold=max(masking threshold, hearing threshold in quiet).
[0165] In one embodiment, the spreading function for a frequency band depends on the sensation level, such that the effect of the spreading function in frequency bands with relatively high sensation levels is greater than the effect of the spreading function in frequency bands with relatively low sensation levels. Typically, the spreading function is defined on an absolute SPL scale. Alternative methods for defining the spreading function may provide a more generalized psychoacoustic model while incurring minimal additional computational complexity. Currently, many encoders appear to apply a spreading function that is most suitable for quiet signals. While this is a conservative design approach, the degree of frequency-domain masking may be underestimated for louder signals. This may lead to allocating more bits than necessary in certain bands, correspondingly reducing the number of bits available for other bands and potentially resulting in reduced quality. Thus, the analysis component 700 may be configured to define the spreading function for that frequency band depending on the difference between the energy value determined in S04 and the hearing threshold in quiet determined in S06, leading to improved bit allocation.
[0166] In some embodiments, determining S18 a bit allocation value for a frequency band includes calculating an SMR for that frequency band, which is the energy value for that frequency band subtracted by the masking threshold calculated S16 for that frequency band. In some embodiments, an additional fixed offset is subtracted. The bit allocation value determination S18 is then based on the amount of SMR. In some embodiments, the bit allocation value is thresholded at a defined maximum bit allocation value, for example, 12 bits.
[0167] In some embodiments, determining S18 the bit allocation value includes adjusting S20 the masking threshold to achieve a bit allocation that meets a target bit rate for the audio signal. Adjusting S20 the masking threshold may include adjusting the masking threshold by adding a constant offset to the masking threshold in the loudness domain until the target bit rate for the audio signal is met. As described above, converting from the linear energy domain to the loudness domain includes raising each energy to approximately the 0.25 to 0.3 power.
[0168] In general, analysis component 700 allocates more bits S18 for frequency bands with higher SV (where SV is defined in dB as a subtractive adjustment to the excitation function) compared to frequency bands with lower SV.
[0169] In some embodiments, the analysis component 700 may be implemented in an encoder 800. Such an embodiment is illustrated in FIG. 8. In this embodiment, the encoder 800 includes a receiving component 802 configured to receive an audio signal 806. The encoder further includes an encoding component 804 configured to use the bit allocation values determined by the analysis component 700 for encoding purposes. For example, the encoding component 804 is configured to quantize audio samples of the audio data for the frequency bands in response to the bit allocation values and encode the quantized audio data for the frequency bands into a bitstream 808. In some embodiments, the encoder 800 further includes a companding component (not shown) configured to reduce the dynamic range of the audio signal using a companding algorithm before quantizing the audio samples of the audio data for the frequency bands. The companding function reduces the dynamic range of the input signal before transform coding. The companding function may be beneficial to the encoded quality of signals that include dense transient mixed sounds, such as rain or applause. In one example, input signal companding and an embodiment that calculates only the first SV S10 and uses this as the SV S08 may work synergistically to produce higher performance than either function separately. In this embodiment, the dynamic range of the audio signal is reduced using a companding algorithm prior to encoding the audio signal using the associated SV. The companding function can further reduce the number of bits to code while still maintaining high audio quality.
[0170] In some embodiments, the analysis component 700 is implemented in a decoder 900. This embodiment is shown in FIG. 9. In this embodiment, the decoder 900 includes a receiving component 902 configured to receive an audio signal 906 in the form of an encoded bitstream that includes encoded energy values for frequency bands of the audio signal. The decoder further includes a decoding component 904 configured to use the bit allocation values determined S18 by the analysis component 700 for decoding purposes. The decoding component 904 is configured to decode the encoded energy values from the encoded bitstream 906, and the analysis component 700 uses the decoded energy values in determining the energy values. The decoding component 904 is further configured to extract quantized audio samples of the audio data for the frequency bands from the encoded bitstream 906 in response to the bit allocation values. The decode component 904 is further configured to dequantize the quantized audio samples of the audio data for the frequency bands and combine the dequantized audio samples of the audio data for each frequency band to generate a decoded audio signal 908.
[0171] It should be noted that the analysis component 700 and corresponding methods can be used with any audio format.
[0172] Masking using the inventive methods described herein provides a better match to subjectively measured masking data, resulting in improved bit allocation. The embodiment using the first calculated SV as the SV for that frequency band provides the greatest improvement over the default encoder for speech signals. This is important because speech signals are a critical component of typical broadcast and movie content.
[0173] In some embodiments, the encoder 800 and / or decoder 900 implementing this embodiment (or alternatively, an embodiment that also calculates S12 the second SV) are implemented in a real-time two-way communication device. Advantageously, this simpler embodiment may be used in such devices due to the lower complexity of such encoding methods. However, it should be noted that there are many applications and possible uses of the encoder 800 and / or decoder 900.
[0174] In this way, the encoder 800 accurately captures the observed masking behavior of the human auditory system, which translates into higher codec performance than the default encoder for both constant bit rate and variable bit rate applications.
[0175] In some embodiments, the subjective improvement may be most noticeable for relatively high-level signals, because default encoders that derive masking thresholds based on level-independent offsets (instead of using SVs) tend to over-allocate bits to low-level signal components.
[0176] FIG. 4 outlines an exemplary method for calculating the masking thresholds described above, in which a tonality estimate is used to weight the first and second SVs. As shown in FIG. 4, an input audio frame is input to an MDCT filter bank. The transform length(s) input to the MDCT filter bank are also received by a tonality estimation unit (discussed further below in connection with FIG. 6). The tonality estimation unit generates a tonality estimate T ranging from 0 to 1, as discussed further below. j (m), one set for each MDCT transform. In this nomenclature, j is the band index and m is the MDCT block index.
[0177] The MDCT transform coefficients are used to determine an energy value for each frequency band. A spreading function is applied to the energy values of the frequency bands to derive an excitation function. In the final steps of the exemplary method of FIG. 4, the energy values and a quiet threshold are used to calculate first and second SVs, which are then used to calculate the tonality estimate T, as described herein. j (m) and applied to the excitation values to finally produce a masking threshold for each frequency band.
[0178] An embodiment of a tonality estimation method based on adaptive prediction will now be described in relation to FIG.
[0179] An input frame 602 provides input samples. A filter bank 604 is configured to receive the input samples from the input frame 602. Note that there are different filter banks 604 that can be used. In one example, a CMDCT is used, with N=128. In another example, a CQMF may be used, with N=64. The filter bank 604 provides complex frequency coefficients 606 (X k 6 is configured to send the unpredictability values 609 (μ (n), band k at time n) corresponding to the frequency coefficients 606 to the LPC analysis component 608 and the unpredictability estimation component 605. The structure of FIG. 6 is repeated for each CMDCT / CQMF band. The LPC analysis component 608, in conjunction with the unpredictability estimation component 605, generates unpredictability values 609 (μ (n), band k at time n) corresponding to the frequency coefficients 606. k (n)).
[0180] Considering the fact that a time sample in one CMDCT block affects three adjacent CMDCT blocks, a 3-tap FIR filter is used to smooth the unpredictability estimate in the two-stage smoothing stage 620. This improves the smoothness of the tonality estimate (and hence the decoded audio). A similar approach is used for other filter banks, for example, CQMF with N=64.
[0181] The mapping component 610 is configured to receive the smoothed unpredictability values, the transform length 612, and the energy of the frequency bands (calculated by box 611 in FIG. 6). After the unpredictability estimates are smoothed, they are combined across time to reflect the non-zero portions of the MDCT window to which they are subsequently applied. This can be important to maximize the clarity of the decoded output signal, especially for dynamically changing signals such as speech. The mapping component 610 further outputs the mapped input data 613 (Z k (n)) as output data to the diffusion and normalization component 614. The diffusion and normalization component 614 applies a diffusion function to the input data 613 to generate a set of corrected data 615 (U k (n)) to a tonality mapping component 616. The tonality mapping component 616 is configured to map the input data 615 (unpredictability) to one or more sets of tonality estimates 618.
[0182] 6, according to some embodiments, a sine window is applied to 50% overlapping blocks of input samples taken from one 4096-long frame 602, followed by a 128-point CMDCT. The choice of filter bank 604 is not critical; for example, a complex QMF filter bank 604 already present in known encoders can be used. A set of complex frequency coefficients 606 X from block 604 is k (n) For k=1,...,N, a corresponding set of unpredictability values 609 is generated. The unpredictability for band k at time n is
number
[0183] At each CMDCT frequency bin k, L k consecutive coefficients X k (nm), m=1,…,L k The group of is windowed to order p k The complex prediction coefficients (p k <L k ) are then analyzed to generate the prediction coefficients a ki , i=1,…,p k Using X k The unpredictability value 609 corresponding to (n) is calculated. Of the various LPC analysis windows evaluated, a nearly symmetric Hamming window was found to maximize the prediction gain. The degree of asymmetry varies as a function of the CMDCT bin. The full-band unpredictability value is then filtered by a set of two-stage smoothing filters 620 to avoid abrupt changes over time. An exemplary two-stage filter consists of a three-tap FIR in cascade with a regular exponential smoothing filter. The FIR filter calculates the unpredictability value 609 μ k (n) and receives the partially smoothed output signal μ k The FIR output signal is then further processed by the exponential smoothing filter to produce μ k "Generate (n).
[0184] In one embodiment, a fast-attack, slow-decay IIR filter may be used for the exponential smoothing filter. These filters provide a means to independently control the attack and decay times. The input is in the tonal domain (1-μ k '(n)), the difference equation is given by:
number
[0185] Without smoothing, tonality estimates tend to fluctuate across successive transform blocks, leading to fluctuations in the masking threshold estimate. This can lead to audible quantization noise modulation in the decoder output, especially at low to mid-frequencies. An effective solution to this problem is to design attack / decay filter coefficients from the known time characteristics of the human auditory filter. This approach generally leads to attack / decay time constants that are longest at low frequencies and shortest at high frequencies.
[0186] In the next stage of Figure 6, the CMDCT bin energies and smoothed unpredictability values from all CMDCT blocks are resampled and combined (mapped) into groups 610 to match the time and frequency resolution of each MDCT transform in the current frame. The resampled unpredictability values 613 are weighted according to their relative perceptual importance and combined across time, if necessary. Exemplary perceptual weightings include the squared L2 norm (energy) and loudness. Next, to match the spreading applied to the banded MDCT energy, the unpredictability values are spread across frequency by applying a spreading function 614, e.g., from ISO / IEC 11172-3:1993(E). In the final step, the resampled, spread, and normalized unpredictability values 615 are mapped to one or more sets of tonality estimates 618 ranging from 0 to 1 (inclusive) (one set for each MDCT transform).
[0187] Figure 10 shows results from an example of experimentally measured SV for JND as a function of various tone-plus-narrowband noise ratios (SNRs). Center frequencies of 500 Hz 804, 1 kHz 802, and 4 kHz 806 were used, with masker SNRs ranging from -10 dB to 40 dB. The masker was presented to the subject at a level of 80 dB SPL. The mean curve 808 (bold line) represents the SV average across all three center frequencies.
[0188] In one embodiment, a different tonality mapping function (different from the tonality mapping rules in ISO / IEC 11172-3:1993) is calibrated based at least in part on the results of a perceptual masking experiment for a mixed tone + narrowband noise signal. The objective of this embodiment is to determine just-in-time noise levels for a masker consisting of a tone + narrowband noise mixture and a maskee consisting of uncorrelated narrowband noise of the same frequency. The experiment is repeated at various masker-tone / noise mixture levels and various frequencies. The results of the experiment can be used to calibrate the tonality mapping, as described below.
[0189] In one embodiment, each tone and narrowband noise stimulus is first injected into a tone estimator to capture the associated unpredictability value. From these results, a table is generated that associates each unpredictability value with the required SMR. This table is combined with a tonality-to-SMR rule for tone and narrowband noise masking that matches the SMR range to derive points on a curve that define the unpredictability-to-tonality mapping required to calibrate the model. In the final step, a parametric function that approximates the derived calibration curve is derived. The masking experiment and calibration steps may be repeated at various frequencies and input signal levels.
[0190] 11 shows, by way of example, experimental results for one embodiment. It presents a target calibration curve 904 and a parametric model / function 902 (dash-dotted line) that approximates the derived calibration curve for a 128-point CMDCT using a sine window and 50% overlap. A prior art example is included in the figure for comparison (dashed line 906). In this example, the LPC analysis is third order with a window length of 6. In this embodiment, the parametric function T(μ) maps unpredictability to tonality as follows: T(μ)=min(max(a*μ 3 +b*μ 2 +c*μ+d,0),1)
[0191] Values for the four parameters (a, b, c, and d) are derived to approximate the target calibration curve. In the example shown in Figure 11, the model parameter values for a, b, c, and d are -13.0233, 15.9513, -8.1012, and 2.1319.
[0192] The use of tonality estimates in perceptual models is well known in the prior art (ISO / IEC 11172-3:1993(E), dashed line 906 in FIG. 11), but the prior art models operate in a level-independent manner. In ISO / IEC 11172-3:1993(E), the SMR model adapts based on estimated tonalities between 6 and approximately 30 dB for pure noise and tonal signals, respectively. The level-dependent model described herein, using a calibrated tonality mapping function, has been realized in simulations and outperforms prior art models as determined by objective quality metrics and listening tests.
[0193] Further embodiments of the present disclosure will be apparent to those skilled in the art after considering the above description. Although the specification and drawings disclose embodiments and examples, the present disclosure is not limited to these particular examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any reference signs appearing in the claims shall not be construed as limiting the scope thereof.
[0194] Furthermore, variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the present disclosure, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0195] The systems and methods disclosed above can be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the division of tasks among the above-described functional units does not necessarily correspond to a division into physical units. Conversely, one physical component may have multiple functions, and one task may be performed by several physical components working together. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those skilled in the art, the term “computer storage media” includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVDs), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and are well known to those skilled in the art to include any information delivery media.
[0196] Various aspects of the present disclosure can be understood from the following enumerated example embodiments (EEE). [EEE1] 1. A method of processing an audio signal, the audio signal comprising audio data in a plurality of frequency bands, the method comprising: For each frequency band of the plurality of frequency bands: Determine the energy value for the audio data in that frequency band; Determine the hearing threshold in quiet for that frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet; calculating a masking threshold for the frequency band using the sensitivity value and the energy value; determining a bit allocation value for the frequency band using the energy value and the masking threshold; method. [EEE2] Calculating the masked threshold comprises: the energy value for the frequency band; or The transformed energy values of the frequency bands applying a spreading function to one of the frequencies to determine an excitation value for that frequency band; combining said sensitivity value with said excitation value; Method described in EEE1. [EEE3] The method according to EEE1, wherein calculating the masking threshold comprises combining the energy value and the sensitivity value to determine an intermediate threshold, and applying a spreading function to the intermediate threshold to determine the masking threshold. [EEE4] 4. The method of any one of EEE1 to 3, further comprising quantizing audio samples of the audio data for the frequency band in response to the bit allocation value. [EEE5] The method according to EEE4, further comprising encoding the quantized audio data of the frequency bands into a bitstream. [EEE6] 4. The method of any one of EEE1 to EEE3, wherein the audio signal is an encoded bitstream comprising encoded energy values for the frequency bands, and wherein determining the energy values for audio data for the frequency bands comprises decoding the encoded energy values from the encoded bitstream. [EEE7] The method according to EEE6, further comprising extracting quantized audio samples of audio data for said frequency bands from said encoded bitstream in response to said bit allocation values. [EEE8] The method of claim 6, further comprising dequantizing the quantized audio samples of the audio data for the frequency bands and combining the dequantized audio samples of the audio data for each frequency band to generate a decoded audio signal. [EEE9] The method of any one of EEE1 to 8, wherein determining the bit allocation value comprises adjusting the masking threshold to achieve a bit allocation that meets a target bit rate for the audio signal. [EEE10] The method according to EEE9, wherein adjusting the masking threshold comprises adjusting the masking threshold by adding a constant offset to the masking threshold in a loudness domain until the target bit rate for the audio signal is met. [EEE11] 11. The method of any one of EEE1 to 10, wherein the energy values, hearing threshold in quiet and masked threshold are defined in decibels dB. [EEE12] 12. The method of any one of EEE1 to 11, further comprising determining the plurality of frequency bands of the audio signal according to an equivalent rectangular bandwidth (ERB) scale. [EEE13] The method of EEE2 or any one of EEE1 to 12 when citing EEE2, wherein the SV is defined in dB as a subtractive adjustment to the excitation value, and wherein determining a bit allocation value comprises allocating more bits for frequency bands with higher SV than for frequency bands with lower SV. [EEE14] 14. The method of any one of EEE1 to 13, wherein the step of calculating an SV for the frequency band comprises calculating a first SV using a sensation level, the sensation level being the difference in dB scale between the energy value and the hearing threshold in quiet. [EEE15] The method of claim EEE14, wherein the step of calculating a first SV comprises multiplying the sensation level by a first scalar. [EEE16] The method of claim 3, wherein the first scalar is frequency dependent. [EEE17] The method of EEE15, wherein the first scalar is constant across all frequency bands. [EEE18] The method of any one of EEE15 to 17, wherein calculating a first SV comprises adding a second scalar to the sensation level multiplied by the first scalar. [EEE19] 19. The method of any one of EEE14 to 18, wherein the step of calculating an SV comprises using the first SV as the SV for the frequency band. [EEE20] 19. The method of any one of EEE14 to 18, wherein the step of calculating an SV for the frequency band comprises calculating a second SV using the perception level, and weighting the first and second SVs based on at least one characteristic of the audio signal. [EEE21] The method of EEE20, wherein calculating a second SV for the frequency band comprises multiplying the sensation level by a third scalar different from the first scalar. [EEE22] The method of claim 8, wherein calculating a second SV comprises adding a fourth scalar to the sensation level multiplied by the third scalar, the fourth scalar being different from the second scalar. [EEE23] 23. The method of any one of EEE20 to 22, wherein weighting the first and second SVs based on at least one characteristic of the audio signal comprises calculating a value representing the weight, the value ranging from 0 to 1, and wherein calculating the SV for the frequency band comprises multiplying one of the first and second SVs by the value, multiplying the other of the first or second SVs by 1 minus the value, and adding the two resulting sums together to form the SV for the frequency band. [EEE24] 24. The method of any one of EEE20 to 23, wherein said at least one characteristic defines an estimated tonality of that frequency band of the audio signal. [EEE25] The method of any one of EEE20 to 23, wherein said at least one characteristic defines an estimated level of noise in that frequency band of the audio signal. [EEE26] The method according to EEE24, wherein the estimated tonality is calculated using adaptive prediction of frequency coefficients calculated from the frequency band of the audio signal. [EEE27] The method according to EEE26, wherein linear predictive coding (LPC) is adaptively applied to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated. [EEE28] The method according to EEE27, wherein the LPC analysis window length is varied as a function of said frequency band. [EEE29] The method according to EEE28, wherein a relatively longer LPC analysis window is used for the relatively lower frequency band. [EEE30] 30. The method according to any one of EEE27 to 29, wherein the LPC prediction order is varied as a function of the frequency band. [EEE31] The method of any one of EEE1 to EEE30, wherein the frequency range of the audio signal is between 200 and 7000 Hz. [EEE32] The method of any one of EEE1 to EEE31, wherein the step of determining the hearing threshold in quiet for the frequency band comprises using a predefined table defining the hearing threshold for at least some frequencies. [EEE33] The method described in EEE4 or any other EEE method when citing EEE4, wherein before quantizing audio samples of the audio data for the frequency bands, the dynamic range of the audio signal is reduced using a companding algorithm. [EEE34] The method of any one of EEE14 or EEE15 to 33 when citing EEE14, further comprising a step of defining a spreading function for a frequency band depending on the sensation level, such that the effect of the spreading function in that frequency band is greater than the effect of the spreading function in that frequency band having a relatively higher sensation level compared to the effect of the spreading function in a frequency band having a relatively lower sensation level. [EEE35] an analysis component configured to receive an audio signal, the audio signal including audio data in a plurality of frequency bands; an analysis component configured to determine a plurality of frequency bands of the audio signal, The analysis component, for each frequency band of the plurality of frequency bands: determining an energy value for said audio data for that frequency band; Determine the hearing threshold in quiet for that frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet; calculating a masking threshold for the frequency band using the sensitivity value and the energy value; and further configured to determine a bit allocation value for the frequency band using the energy value and the masking threshold. Device. [EEE36] The analysis component: an energy value for said frequency band; or The transformed energy values of the frequency bands determining an excitation value for the frequency band by applying a spreading function to one of the By combining the sensitivity value with the excitation value, 8. The apparatus of claim 35, configured to calculate the masking threshold. [EEE37] 3. The apparatus of claim 8, wherein the analysis component is configured to calculate the masking threshold by combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold. [EEE38] 38. The apparatus of any one of claims 35 to 37, being an encoder, further comprising an encoding component configured to quantize audio samples of audio data for the frequency bands in response to the bit allocation values. [EEE39] The apparatus of EEE38, wherein the encoding component is further configured to encode the quantized audio data for the frequency bands into a bitstream. [EEE40] 10. The apparatus of any one of claims 8 to 9, wherein the audio signal is an encoded bitstream including encoded energy values for the frequency bands, the apparatus further comprising a decoding component configured to decode the encoded energy values from the encoded bitstream, and the analysis component uses the decoded energy values when determining the energy values. [EEE41] The apparatus of EEE40, wherein the decoding component is further configured to extract quantized audio samples of audio data for the frequency bands from the encoded bitstream in response to the bit allocation value. [EEE42] The apparatus of EEE41, wherein the decoding component is further configured to dequantize quantized audio samples of the audio data for the frequency bands and combine the dequantized audio samples of the audio data for each frequency band to generate a decoded audio signal. [EEE43] 43. The apparatus of any one of EEE35 to 42, wherein the analysis component, when determining the bit allocation value, is configured to adjust the masking threshold to achieve a bit allocation that meets a target bit rate for the audio signal. [EEE44] The apparatus of EEE43, wherein the analysis component is configured to adjust the masking threshold by adding a constant offset to the masking threshold in a loudness domain until the target bit rate for the audio signal is met. [EEE45] 45. The apparatus of any one of EEE35 to 44, wherein the analysis component is configured to define the energy values, hearing thresholds in quiet and masking thresholds in decibels dB. [EEE46] 6. The apparatus of any one of EEE35 to 45, wherein the analysis component is configured to determine the plurality of frequency bands of the audio signal according to an equivalent rectangular bandwidth (ERB) scale. [EEE47] 5. The apparatus of any one of EEE36 or EEE37-46, wherein the SV is defined in dB as a subtractive adjustment to the excitation value, and wherein the analysis component is configured to determine a bit allocation value by allocating more bits for frequency bands with higher SV than for frequency bands with lower SV. [EEE48] 8. The apparatus of any one of EEE35 to 47, wherein the analysis component is configured to calculate a SV for the frequency band by calculating a first SV using a sensation level, the sensation level being the difference in dB scale between the energy value and the hearing threshold in quiet. [EEE49] 8. The apparatus of claim 8, wherein the analysis component is configured to calculate the first SV by multiplying the sensation level by a first scalar. [EEE50] The apparatus of EEE49, wherein the first scalar is frequency dependent. [EEE51] 8. The apparatus of claim 7, wherein the first scalar is constant across all frequency bands. [EEE52] 52. The apparatus of any one of EEE49 to 51, wherein the analysis component is configured to calculate the first SV by adding a second scalar to the sensation level multiplied by the first scalar. [EEE53] 53. The apparatus of any one of EEE48 to 52, wherein the analysis component is configured to calculate the SV by using the first SV as the SV for that frequency band. [EEE54] 53. The apparatus of any one of EEE48 to 52, wherein the analysis component is configured to further calculate a second SV using the sensation level, and to calculate an SV for the frequency band by weighting the first and second SV based on at least one characteristic of the audio signal. [EEE55] The apparatus of EEE54, wherein the analysis component is configured to calculate the second SV for the frequency band by multiplying the sensation level by a third scalar different from the first scalar. [EEE56] 8. The apparatus of claim 6, wherein the analysis component is configured to calculate the second SV by adding a fourth scalar to the sensation level multiplied by the third scalar, the fourth scalar being different from the second scalar. [EEE57] 56. The apparatus of any one of claims EEE54 to 55, wherein the analysis component is configured to weight the first and second SVs based on at least one characteristic of the audio signal by calculating a value representing the weight, the value ranging from 0 to 1; multiplying one of the first and second SVs by the value and multiplying the other of the first or second SVs by 1 minus the value; and adding the two resulting sums together to form the SV for the frequency band. [EEE58] 58. The apparatus of any one of EEE54 to 57, wherein the at least one characteristic defines an estimated tonality of that frequency band of the audio signal. [EEE59] 58. The apparatus of any one of EEE54 to 57, wherein the at least one characteristic defines an estimated level of noise in that frequency band of the audio signal. [EEE60] The apparatus of EEE58, wherein the analysis component is configured to calculate the estimated tonality using adaptive prediction of frequency coefficients calculated from the frequency band of the audio signal. [EEE61] The apparatus of EEE60, wherein the analysis component is configured to adaptively apply LPC to MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated. [EEE62] 8. The apparatus of claim 61, wherein an LPC analysis window length is varied as a function of the frequency band. [EEE63] The apparatus described in EEE62, wherein a relatively longer LPC analysis window is used for a relatively lower frequency band. [EEE64] 64. Apparatus according to any one of EEE62 to 63, wherein the LPC prediction order is varied as a function of the frequency band. [EEE65] 10. The device according to any one of EEE35 to 64, wherein the frequency range of the audio signal is between 200 and 7000 Hz. [EEE66] 10. The apparatus of any one of EEE35 to 65, further comprising a memory storing a table defining hearing thresholds in quiet for at least some frequencies, wherein the analysis component is configured to determine the hearing thresholds in quiet for the frequency bands by using the predefined table. [EEE67] The apparatus of EEE38 or any other EEE when citing EEE38, further comprising a companding component configured to reduce a dynamic range of the audio signal using a companding algorithm before quantizing audio samples of the audio data for the frequency bands. [EEE68] 68. The apparatus of any one of EEE48 or EEE49 to 67, wherein the analysis component is configured to define spreading functions for the frequency bands in dependence on the sensation level such that an effect of the spreading function in frequency bands having a relatively higher sensation level is greater compared to an effect of the spreading function in frequency bands having a relatively lower sensation level. [EEE69] 10. The apparatus of any one of claims 8 to 9, implemented in a real-time two-way communication device. [EEE70] 1. A method for estimating tonality of an input signal, comprising: applying a filter bank to achieve a set of frequency coefficients; and calculating an estimated tonality using adaptive prediction of the frequency coefficients. method. [EEE71] The method according to EEE70, wherein the step of calculating estimated tonality comprises applying adaptive linear prediction to said frequency coefficients based on a frequency band of the audio signal from which said frequency coefficients are calculated. [EEE72] The method of EEE70, wherein the LPC analysis window length is varied as a function of said frequency band. [EEE73] The method according to EEE72, in which a relatively longer LPC analysis window is used for the relatively lower frequency band. [EEE74] 74. The method of any one of EEE72 to 73, wherein the LPC prediction order is varied as a function of the frequency band. [EEE75] The method of any one of EEE70 to EEE74, wherein the filter bank comprises one of a 128-band complex MDCT or DFT filter bank and a 64-band complex QMF filter bank. [EEE76] The method according to any one of EEE71 to 73, wherein said LPC analysis window is an asymmetric Hamming window. [EEE77] 77. The method of any one of EEE70 to 76, comprising weighting predictability measures from the adaptive prediction according to the relative perceptual importance of each predictability measure. [EEE78] The method according to EEE77, wherein the step of weighting the predictability measures contained within each time-frequency tile comprises one of weighting based on the energy or loudness of the input signal. [EEE79] The method of any one of EEE70 to 78, further comprising combining a predictability measure from the adaptive prediction of the frequency coefficients to match the time and frequency resolution of the filter bank. [EEE80] The method of any one of EEEE2 or EEE4 to 34 when citing EEE2, wherein the sensitivity value and the excitation value are defined in decibels dB and the combining step comprises subtracting the sensitivity value from the excitation value, or wherein the sensitivity value and the excitation value are defined on an intensity scale and the combining step comprises calculating a quotient of the excitation value and the sensitivity value. [EEE81] The method of EEEE3 or any one of EEE4 to 34 when citing EEE3, wherein the energy value and the sensitivity value are defined in decibels dB and the combining step comprises subtracting the sensitivity value from the energy value, or wherein the energy value and the sensitivity value are defined on an intensity scale and the combining step comprises calculating a quotient of the energy value and the sensitivity value. [EEE82] The method of any one of EEE1 to EEE34 or EEE80 to EEE81, wherein calculating the sensitivity value comprises calculating a ratio or difference between the energy value of the frequency band and the hearing threshold in quiet for the frequency band. [EEE83] A computer program product comprising a computer readable storage medium having instructions adapted to perform the method of any one of EEE1 to EEE34 or EEE80 to EEE82 when executed by a device having processing capability. [EEE84] A computer program product comprising a computer readable storage medium having instructions adapted to perform the method of any one of claims EEE70 to EEE70 when executed by a device having processing capability. [EEE85] The apparatus of EEE36, or any one of EEE38 to EEE69 when citing EEEE36, wherein the sensitivity value and the excitation value are defined in decibels dB and the combining step comprises subtracting the sensitivity value from the excitation value, or wherein the sensitivity value and the excitation value are defined on an intensity scale and the combining step comprises calculating a quotient of the excitation value and the sensitivity value. [EEE86] 8. The apparatus of claim EEE37, or any one of EEE38 to EEE69 when citing EEE37, wherein the energy value and the sensitivity value are defined in decibels dB and the combining step comprises subtracting the sensitivity value from the energy value, or wherein the energy value and the sensitivity value are defined on an intensity scale and the combining step comprises calculating a quotient of the energy value and the sensitivity value. Red text is for division
Claims
1. 1. A method for encoding an audio signal, executed by one or more processors, the method comprising: receiving an input frame of the audio signal; converting the input frames of the audio signal into a frequency domain audio signal comprising audio data in a plurality of frequency bands; For each frequency band of the plurality of frequency bands: determining an energy value for the audio data in the frequency band; determining a hearing threshold in quiet for the frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet, wherein calculating the sensitivity value comprises calculating a ratio or difference between the energy value for the frequency band and the hearing threshold in quiet for that frequency band; calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes: applying a spreading function to one of the energy value for the frequency band; or the transformed energy value of the frequency band to determine an excitation value for the frequency band; and combining the sensitivity value with the excitation value; determining a bit allocation value for the frequency band using the energy value and the masking threshold; quantizing audio samples of the audio data for the frequency band in response to the bit allocation values; encoding the quantized audio data of the frequency bands into a bitstream; including performing method.
2. 1. A method for encoding an audio signal, executed by one or more processors, the method comprising: receiving an input frame of the audio signal; converting the input frames of the audio signal into a frequency domain audio signal comprising audio data in a plurality of frequency bands; For each frequency band of the plurality of frequency bands: determining an energy value for the audio data in the frequency band; determining a hearing threshold in quiet for the frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet, wherein calculating the sensitivity value comprises calculating a ratio or difference between the energy value for the frequency band and the hearing threshold in quiet for that frequency band; calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold; determining a bit allocation value for the frequency band using the energy value and the masking threshold; quantizing audio samples of the audio data for the frequency band in response to the bit allocation values; encoding the quantized audio data of the frequency bands into a bitstream; Including, method.
3. 3. The method of claim 1, wherein determining a bit allocation value comprises allocating more bits for frequency bands with higher SVs than for frequency bands with lower SVs.
4. 3. The method of claim 1, wherein calculating an SV for the frequency band comprises calculating a first SV using a sensation level, the sensation level being the difference in dB scale between the energy value and the hearing threshold in quiet.
5. 5. The method of claim 4, wherein calculating a first SV includes multiplying the sensation level by a first scalar, and / or calculating an SV includes using the first SV as the SV for the frequency band.
6. 5. The method of claim 4, wherein calculating an SV for the frequency band comprises calculating a second SV using the sensation level and weighting the first and second SVs based on at least one characteristic of the audio signal.
7. The method of claim 6 , wherein the at least one characteristic defines an estimated level of tonality in that frequency band of the audio signal.
8. 8. The method of claim 7, wherein the estimated level of tonality is calculated using adaptive prediction of frequency coefficients calculated from that frequency band of the audio signal.
9. 9. The method of claim 8, wherein linear predictive coding (LPC) is adaptively applied to the MDCT coefficients based on the frequency band of the audio signal from which the MDCT coefficients are calculated.
10. 10. The method of claim 9, wherein an LPC analysis window length is varied as a function of the frequency band and / or a prediction order of the LPC is varied as a function of the frequency band.
11. 5. The method of claim 4, wherein the spreading function for the frequency band depends on the sensation level such that the effect of the spreading function in frequency bands having relatively higher sensation levels is greater compared to the effect of the spreading function in frequency bands having relatively lower sensation levels.
12. 12. The method of claim 11, wherein before quantizing audio samples of the audio data for the frequency bands, the dynamic range of the audio signal is reduced using a companding algorithm.
13. a receiving component configured to receive input frames of an audio signal; a filterbank configured to transform the input frames of the audio signal into a frequency domain audio signal comprising audio data in a plurality of frequency bands; an analysis component, The analysis component, for each frequency band of the plurality of frequency bands: determining an energy value for the audio data in that frequency band; determining a hearing threshold in quiet for the frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet, wherein calculating the sensitivity value comprises calculating a ratio or difference between the energy value for the frequency band and the hearing threshold in quiet for that frequency band; calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes: applying a spreading function to one of the energy value for the frequency band; or the transformed energy value of the frequency band to determine an excitation value for the frequency band; and combining the sensitivity value with the excitation value; an analysis component configured to perform a step of determining bit allocation values for the frequency bands using the energy values and the masking thresholds; For each frequency band of the plurality of frequency bands: quantizing audio samples of the audio data for the frequency bands in response to the bit allocation values; an encoding component configured to encode the quantized audio data of the frequency bands into a bitstream; Encoding device.
14. a receiving component configured to receive input frames of an audio signal; a filterbank configured to transform the input frames of the audio signal into a frequency domain audio signal comprising audio data in a plurality of frequency bands; an analysis component, The analysis component, for each frequency band of the plurality of frequency bands: determining an energy value for the audio data in that frequency band; determining a hearing threshold in quiet for the frequency band; calculating a sensitivity value (SV) for the frequency band using the energy value and the hearing threshold in quiet, wherein calculating the sensitivity value comprises calculating a ratio or difference between the energy value for the frequency band and the hearing threshold in quiet for that frequency band; calculating a masking threshold for the frequency band using the sensitivity value and the energy value, wherein calculating the masking threshold includes combining the energy value and the sensitivity value to determine an intermediate threshold and applying a spreading function to the intermediate threshold to determine the masking threshold; an analysis component configured to perform a step of determining bit allocation values for the frequency bands using the energy values and the masking thresholds; For each frequency band of the plurality of frequency bands: quantizing audio samples of the audio data for the frequency bands in response to the bit allocation values; an encoding component configured to encode the quantized audio data of the frequency bands into a bitstream; Encoding device.
15. A computer program product having a computer readable storage medium having instructions adapted to perform the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Method and device for allocating dynamic bit for audio coding
JP2000004163A