Companding apparatus and method to reduce quantization noise using advanced spectral extension

The method addresses quantization noise in audio codecs by dividing audio signals into short segments, applying non-energy-based gains, and using filterbanks to smooth gain application, effectively reducing noise and maintaining audio quality.

JP2025160357APending Publication Date: 2025-10-22DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025126002
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-09-12
Filing Date
2025-07-29
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Current audio codecs introduce noticeable quantization noise due to lossy compression, particularly during low-intensity segments, leading to pre-echo distortion and audible artifacts like clicks, while existing solutions like filters cause phase distortion and reduced frequency resolution.

Method used

A method involving a compression process that divides audio signals into short segments, applies non-energy-based wideband gains to amplify low-intensity segments and attenuate high-intensity segments, followed by an expansion process to restore the original dynamic range, using filterbanks like QMF or STFT to smooth gain application and minimize discontinuities.

Benefits of technology

Effectively reduces quantization noise by shaping it to follow the temporal envelope of the original signal, making it less audible during quiet passages and maintaining audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160357000001_ABST
    Figure 2025160357000001_ABST
Patent Text Reader

Abstract

To provide a companding method and system to reduce encoding noise in an audio codec.SOLUTION: A compression process includes: reducing a dynamic range of an audio signal; dividing the audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain for each time segment in a frequency domain using a non-energy based average of a frequency domain representation of the audio signal; and applying individual gain values so as to amplify segments of relatively low intensity and attenuate segments of relatively high intensity. Next, the compressed audio signal is expanded to return to a substantially original dynamic range. Inverse gain values are applied to amplify the segments of relatively high intensity and attenuate the segments of relatively low intensity. A QMF filter bank is used to obtain a frequency domain representation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One or more embodiments of the present invention relate generally to audio signal processing, and more particularly to reducing coding noise in audio codecs using compression / expansion techniques.

[0002] This application claims priority to U.S. Provisional Patent Application No. 61 / 809,028, filed April 5, 2013, and U.S. Provisional Patent Application No. 61 / 877,167, filed September 12, 2013, all of which are incorporated herein by reference. [Background technology]

[0003] Many common digital sound formats use lossy data compression techniques, discarding some data to reduce storage or due to data rate requirements. The application of lossy data compression not only reduces the fidelity of the source content (e.g., audio content), but also introduces noticeable distortions in the form of compression artifacts. In the context of audio coding systems, these sound artifacts are called coding noise or quantization noise.

[0004] Digital audio systems use codecs (encoding-decoding components) to compress and decompress audio data according to defined audio file formats or streaming media audio formats. Codecs implement algorithms that attempt to represent audio signals using the minimum number of bits while maintaining as high a fidelity as possible. Lossy compression techniques typically used in audio codecs operate on a psychoacoustic model of human hearing. Audio formats often involve the use of time-frequency transforms (e.g., modified discrete cosine transform - MDCT) and employ masking effects, such as frequency masking or temporal masking, where a given sound, including any apparent quantization noise, is hidden or masked by the actual content.

[0005] Most audio coding is frame-based. Within a frame, audio codecs typically shape coding noise in the frequency domain to minimize its audibility. Some current digital audio formats use frames of long duration so that a frame can contain sounds of several different levels or intensities. Because coding noise usually does not vary in level over the course of a frame, coding noise is most audible during portions of the frame where intensity is low. Such an effect can be described as pre-echo distortion, where silence (or low-level signals) preceding a high-intensity segment is immersed in noise in the decoded audio signal. Such an effect can be most noticeable in transient sounds or impulses from percussion instruments, such as castanets or other sharp-hitting sources. Such distortion is typically caused by quantization noise introduced in the frequency domain, which is spread across the codec's transform window in the time domain.

[0006] Current means for avoiding or minimizing pre-echo artifacts include the use of filters. Such filters, however, introduce phase distortion and temporal smearing. Another possible solution involves the use of smaller transform windows. However, this approach can significantly reduce frequency resolution.

[0007] The technical matters described in the Background section should not be assumed to be prior art merely as a result of their mention in the Background section. Similarly, the problems mentioned in the Background section or problems related to the technical matters in the Background section should not be assumed to have been previously recognized in the prior art. The technical matters in the Background section merely illustrate different approaches and may be inventions in themselves. Summary of the Invention [Problem to be solved by the invention]

[0008] An embodiment of the present invention is directed to a method for processing a received audio signal by expanding the audio signal to an extended dynamic range through a process including dividing the received audio signal into a plurality of time segments using a defined window shape, calculating a wideband gain for each time segment in the frequency domain using a non-energy-based average of a frequency-domain representation of the audio signal, and applying a gain value to each time segment to obtain the expanded audio signal. The wideband gain value applied to each time segment is selected to have the effect of amplifying relatively high-intensity segments and attenuating relatively low-intensity segments. For this method, the received audio signal includes an original audio signal that has been compressed from its original dynamic range through a compression process including dividing the original audio signal into a plurality of time segments using a defined window shape, calculating a wideband gain in the frequency domain using a non-energy-based average of frequency-domain samples of the original audio signal, and applying the wideband gain to the original audio signal. In the compression process, the gain value of the wideband gain applied to each time segment is selected to have the effect of amplifying segments of relatively low intensity and attenuating segments of relatively high intensity. The expansion process is configured to substantially restore the dynamic range of the original audio signal, and the wideband gain of the expansion process may be substantially the inverse of the wideband gain of the compression process. [Means for solving the problem]

[0009] In a system implementing a method for processing a received audio signal through an enhancement process, a filterbank component may be used to analyze the audio signal to obtain a frequency-domain representation, and a window shape defined for the division into multiple time segments may be identical to a prototype filter for the filterbank. Similarly, in a system implementing a method for processing a received audio signal through a compression process, a filterbank component may be used to analyze the original audio signal to obtain a frequency-domain representation, and a window shape defined for the division into multiple time segments may be identical to a prototype filter for the filterbank. In either case, the filterbank may be a QMF bank or a short-time Fourier transform. In this system, the received signal for the enhancement process is obtained after transformation of the compressed signal by an audio encoder that generates a bitstream and a decoder that decodes the bitstream. The encoder and decoder may include at least part of a transform-based audio codec. The system may further include a component that processes control information received through the bitstream and that determines the activation state of the enhancement process. [Brief explanation of the drawings]

[0010] In the following drawings, like reference numerals are used to refer to like elements. The following drawings illustrate various embodiments, but one or more implementations are not limited to the embodiments shown in the drawings. [Figure 1] FIG. 1 illustrates a system for compressing and expanding an audio signal in a transform-based audio codec, under one embodiment. [Figure 2A] FIG. 2A shows an audio signal divided into multiple short-time segments, under one embodiment. [Figure 2B]FIG. 2B shows the audio signal of FIG. 2A after application of a wideband gain over each short-time segment, under one embodiment. [Figure 3A] FIG. 3A is a flow chart illustrating a method for compressing an audio signal, under one embodiment. [Figure 3B] FIG. 3B is a flowchart illustrating a method for enhancing an audio signal, under one embodiment. [Figure 4] FIG. 4 is a block diagram illustrating a system for compressing an audio signal, under one embodiment. [Figure 5] FIG. 5 is a block diagram illustrating a system for enhancing an audio signal, under one embodiment. [Figure 6] FIG. 6 illustrates the division of an audio signal into multiple short-time segments, under one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] The use of companding techniques to achieve temporal shaping of quantization noise in audio codecs is described. Such embodiments include the use of companding algorithms implemented in the QMF domain to achieve temporal shaping of quantization noise. The process includes encoder control of the desired decoder companding level, and includes extension beyond mono applications to stereo and multi-channel companding.

[0012] Aspects of one or more embodiments described herein may be implemented in an audio system that processes audio signals for transmission over a network, including one or more computers or processing devices executing software instructions. The described embodiments may be used alone or with other embodiments in any combination. Various embodiments have been motivated by various deficiencies in the prior art, and although described or referenced in one or more places herein, an embodiment need not address every single one of these deficiencies. In other words, different embodiments may address different deficiencies described herein. Some embodiments may only partially address some deficiencies or address only one deficiency described herein. And some embodiments may not address any of these deficiencies at all.

[0013] FIG. 1 illustrates a compression / decompression system for reducing quantization noise in a codec-based audio processing system, under one embodiment. FIG. 1 illustrates an audio signal processing system built around an audio codec, including an encoder (or "core encoder") 106 and a decoder (or "core decoder") 112. The encoder 106 encodes audio content into a data stream or signal for transmission over a network 110, where it is decoded by a decoder 112 for playback or further processing. In one embodiment, the codec's encoder 106 and decoder 112 implement a lossy compression method to reduce the storage and / or data rate of digital audio data. Such a codec may be implemented as an MP3, Vorbis, Dolby Digital (AC-3), AAC, or similar codec. The codec's lossy compression method generates coding noise, which generally has no variation in level across the frame spread defined by the codec. Such coding noise is often most audible during low-intensity portions of the frame. The system 100 includes components that reduce the perceived coding noise in existing coding systems by providing a compression pre-step component 104 prior to the codec's core encoder 106 and an enhanced post-step component 114 that operates on the core decoder 112 output. The compression component 104 is configured to divide the original audio input signal 102 into multiple time segments using a defined window shape, compute a wideband gain in the frequency domain using a non-energy-based average of frequency-domain samples of the original audio signal, and apply a wideband gain in the frequency domain using a non-energy-based average of frequency-domain samples of the original audio signal. Here, the gain value applied to each time segment amplifies segments with relatively low intensity and attenuates segments with relatively high intensity. This gain modification has the effect of compressing, or significantly reducing, the original dynamic range of the input audio signal 102.The compressed audio signal is then encoded in an encoder 106, transmitted over a network 110, and decoded in a decoder 112. The decoded compressed signal is input to an expansion component 114. The expansion component is configured to perform the inverse operation of the compression pre-step 104, expanding the dynamic range of the compressed audio signal back to that of the original input audio signal 102 by applying an inverse gain value to each time segment. Thus, the audio output signal 116 contains an audio signal with its original dynamic range, with coding noise removed through the pre-step and post-step compression and expansion processes.

[0014] As shown in FIG. 1, the compression component or compression pre-step 104 is configured to reduce the dynamic range of the audio signal 102 input to the core encoder 106. The input audio signal is divided into a number of short segments. The size or length of the short segments is a fraction of the frame size used by the core encoder 106. For example, a typical frame size for a core coder may be on the order of 40 to 80 milliseconds. In this case, each short segment may be on the order of 1 to 3 milliseconds. The compression component 104 calculates appropriate wideband gain values ​​for compressing the input audio signal for each segment. This is achieved by modifying the short segments of the signal with an appropriate gain value for each segment. Relatively large gain values ​​are selected to amplify segments with relatively low intensity, and small gain values ​​are selected to attenuate segments with high intensity.

[0015] FIG. 2A illustrates an audio signal divided into multiple short-time segments, and FIG. 2B illustrates the same audio signal after application of wideband gain by a compression component, under one embodiment. As shown in FIG. 2A, audio signal 202 represents a transient or impulse sound, such as that produced by a percussion instrument (e.g., castanets). The signal is characterized by spikes in amplitude, as shown in a plot of voltage V versus time t. Generally, the amplitude of a signal relates to acoustic energy or sound intensity and represents a measure of the sound's power at any instant in time. When audio signal 202 is processed through a frame-based audio codec, signal portions are processed within a transform (e.g., MDCT) frame 204. Typical current audio systems use frames of relatively long duration, such that for sharp transient or short impulse sounds, a single frame can contain both high-intensity and low-intensity sounds. Thus, as shown in FIG. 1 , one MDCT frame 204 contains the impulse portion (peak) of the audio signal, as well as a relatively large amount of low-intensity signal before and after the peak. In one embodiment, the compression component 104 divides the signal into a number of short segments 206 and applies a wideband gain to each segment to compress the dynamic range of the signal 202. The number and size of each short segment can be selected based on application needs and system limitations. With respect to the size of an individual MDCT frame, the number of short segments can be between 12 and 64 segments, with 32 segments being typical. However, the embodiment is not so limited.

[0016] Figure 2B shows the audio signal of Figure 2A after application of a wideband gain across each short-time segment, under one embodiment. As shown in Figure 2B, audio signal 212 has relatively the same shape as original signal 202. However, the amplitude of low-intensity segments has been increased by application of an amplifying gain value, and the amplitude of high-intensity segments has been decreased by application of an attenuating gain value.

[0017] The output of the core decoder 112 is the input audio signal (e.g., signal 212) with reduced dynamic range and the quantization noise introduced by the core encoder 106. This quantization noise is characterized by a nearly uniform level across time within each frame. The enhancement component 114 operates on the decoded signal to restore the dynamic range of the original signal. The enhancement component uses the same short temporal resolution based on the short segment size 206 and inverts the gain applied in the compression component 104. Thus, the enhancement component 114 applies a small gain (attenuation) to segments that were low intensity in the original signal, which would have been amplified by the compressor, and a large gain (amplification) to segments that were high intensity in the original signal, which would have been attenuated by the compressor. The quantization noise added by the core coder had a uniform temporal envelope, but is thus simultaneously shaped by the post-processor gain to approximately follow the temporal envelope of the original signal. This process effectively makes the quantization noise less audible during quiet passages. The noise may be amplified during high intensity passages, but remains less audible due to the masking effect of the loud signal of the audio content itself.

[0018] As shown in FIG. 2A , the compression / expansion process modifies individual segments of an audio signal individually with respective gain values. In certain cases, this can result in discontinuities in the output of the compression component, which can cause problems in the core encoder 106. Similarly, gain discontinuities in the expansion component 114 can result in discontinuities in the envelope of the generated noise. The noise can result in audible clicks in the audio output 116. Another problem with applying individual gain values ​​to short segments of an audio signal stems from the fact that a typical audio signal is a combination of many individual sources. Some of these sources may be stationary over time, and some may be transient. Stationary signals are generally constant over time in statistical parameters, while transient signals are generally not. Given the wideband nature of transients, fingerprints in such mixtures are often more visible at higher frequencies. Gain calculations based on the short-term energy (RMS) of the signal tend to be biased to be stronger at low frequencies, where stationary sources dominate and show little change over time. This energy-based approach is therefore generally ineffective at shaping the noise introduced by the core encoder.

[0019] In one embodiment, the system 100 calculates and applies gains in the compression and expansion components in a filter-bank using short prototype filters to solve potential problems associated with applying individual gain values. The signal to be modified (the original signal in the compression component 104 and the output of the core decoder 112 in the expansion component 114) is first analyzed by the filter-bank, and wideband gains are applied directly in the frequency domain. The corresponding effect in the time domain is to naturally smooth the gain application according to the shape of the prototype filter. This solves the discontinuity problem mentioned above. The modified frequency-domain signal is then transformed back to the time domain through a corresponding unified filter-bank. Analysis of the signal using the filter-bank provides access to the spectral content, allowing gain calculations to preferentially boost high-frequency contributions (or to boost contributions from any weak spectral content), providing gain values ​​that are not dominated by the strongest components in the signal. This solves the problem, as discussed above, for audio sources that contain a combination of different sources. In one embodiment, the system calculates the gain using p-norm spectral magnitude, where p is typically less than two (p<2), which allows for more emphasis on weaker spectral components compared to an energy (p=2) based approach.

[0020] As mentioned above, the system includes a prototype filter to smooth the gain application. Typically, the prototype filter is a basic window shape in a filter bank, modulated by a sinusoidal waveform to obtain the impulse response for the different subband filters in the filter bank. For example, a short-time Fourier transform (STFT) is a filter bank, and each frequency line of this transform is a subband of the filter bank. The short-time Fourier transform is performed by multiplying the signal by a window shape (an N-sample window). The window shape can be rectangular, Hann, Kaiser-Bessel-derived (KBD), or several other shapes. The windowed signal is then subjected to a discrete Fourier transform (DFT) operation to obtain the STFT. The window shape in this case is the prototype filter. The DFT consists of sinusoidal-based functions, each at a different frequency. The window shape multiplied by the sinusoidal function then provides the filter for the subband corresponding to that frequency. It is referred to as a "prototype" because the window shape is the same at all frequencies.

[0021] In one embodiment, the system uses a QMF (Quadrature Modulated Filter) bank as the filter bank. In certain embodiments, the QMF bank may have a 64-point window to form a filter type. This window (corresponding to 64 evenly spaced frequencies) modulated by cosine and sine functions forms a subband filter for the QMF bank. After each application of the QMF function, the window is shifted by 64 samples. This means that the overlap between time segments in this case is 640 - 64 = 576 samples. However, although the window shape in this case spans 10 time segments (640 = 10 * 64), the main lobe of the window (where the sample values ​​are most significant) is approximately 128 samples long. Thus, the effective length of the window is still relatively short.

[0022] In one embodiment, the enhancement component 114 ideally inverts the gain applied by the compression component 104. While the gain applied by the compression component could be transmitted to the decoder via a bitstream, such an approach typically consumes significant bitrate. In one embodiment, the system 100 instead estimates the gain required by the enhancement component 114 directly from the available signal, i.e., the output of the decoder 112. This is efficient and does not require additional bits. The filter banks in the compression and enhancement components are chosen to be identical in order to compute gains that are the inverse of each other. Additionally, these filter banks are time-synchronized, so that any effective delay between the output of the compression component 104 and the input to the enhancement component 114 is a multiple of the filter bank stride. If the core encoder-decoder were lossless and the filter banks provided perfect recovery, the gains in the compression and enhancement components would be the exact inverse of each other. In this way, exact recovery of the original signal is achieved. In practice, however, the gain provided by the expansion component 114 is only a very close approximation of the inverse of the gain applied by the compression component 104 .

[0023] In one embodiment, the filter bank used in the compression and expansion components is a QMF bank. In a typical application, a core audio frame may be 4096 samples long, with a 2048 overlap with adjacent frames. At 48 kHz, such a frame is 85.3 milliseconds long. In contrast, the QMF bank used may have a stride of 64 samples (1.3 milliseconds long), providing fine temporal resolution for gain. Furthermore, the QMF has a smooth prototype filter 640 samples long, ensuring that gain application varies smoothly over time. Analysis using this QMF filter bank provides a time-frequency tiled representation of the signal. Each QMF time slot is equal to the stride, and in each QMF time slot, there are 64 uniformly spaced subbands. Alternatively, other filter banks can be used, such as short-time Fourier transforms (STFTs), and such a time-frequency tiled representation can still be obtained.

[0024] In one embodiment, the compression component 104 performs a pre-processing step to condition the codec input. For this embodiment, S t (k) is a complex-valued filterbank sample at time slot t and frequency bin k. Figure 6 illustrates the division of an audio signal into a number of time slots for a range of frequencies, under one embodiment. For the example of diagram 600, there are 64 frequency bins k and 32 time slots t, which generate multiple time-frequency tiles as shown (not necessarily to scale). The compression pre-step is performed by dividing the codec input into S' t (k)=S t (kg t In this equation,

number

[0025] In the above equation,

number

[0026]

number

[0027] The 1-norm has been shown to give significantly better results than using the energy (rms / 2-norm). The value of the exponential term γ typically ranges between 0 and 1 and may be chosen to be 1 / 3. The constant S0 ensures reasonable gain values ​​independent of the implementation platform. For example, for all S t If the (k) value is implemented on a platform where the absolute value is limited to 1, the constant S0 may be 1. t It can potentially be different on platforms where (k) may have a different maximum value. It can also be used to ensure that the average gain value over a large set of signals is close to 1; that is, it can be an intermediate signal value between the maximum and minimum signal values ​​determined from a large collection of content.

[0028] In a post-step process performed by the expansion component 114, the output is expanded by the inverse gain applied by the compression component 104. This requires an exact or near-exact replica of the filter bank of the compression component. In this case,

number

number

[0029] In the above equation,

number

[0030]

number

number

[0031] Generally, the expansion component 114 uses the same p-norm as that used in the compression component 104. Thus,

number

number

[0032] When complex filter banks (containing both cosine and sine based functions), such as STFT or complex-QMF, are used in the compression and expansion components, Magnitude, i.e., the magnitude of the decoded subband samples

number

number

[0033] In the above equation, the value K is less than or equal to the number of subbands in the filterbank. In general, the p-norm can be calculated using any subset of subbands in the filterbank. However, the same subset should be used in both the encoder 106 and the decoder 112. In one embodiment, the high-frequency portion of the audio signal (e.g., audio components above 6 kHz) can be coded using advanced spectrum extension (A-SPX) tools. Additionally, it may be desirable to use only signals above 1 kHz (or similar frequencies) to guide noise shaping. In such cases, only those subbands in the 1 kHz to 6 kHz range can be used to calculate the p-norm, and therefore the gain value. Furthermore, although the gain is calculated from one subset of subbands, it can be applied to different, and possibly even larger, subsets of subbands.

[0034] As shown in Figure 1, the compression function of shaping the quantization noise introduced by the core encoder 106 of the audio codec is performed by two separate components 104 and 114 that perform certain pre-encoder compression functions and post-decoder extension functions. Figure 3A is a flowchart illustrating a method for compressing an audio signal in the pre-encoder compression component, and Figure 3B is a flowchart illustrating a method for extending an audio signal in the post-decoder extension component, under one embodiment.

[0035] As shown in FIG. 3A, process 300 begins with a compression component receiving an input audio signal (302). The compression then divides the audio signal into short-time segments (304) and compresses the audio signal to a reduced dynamic range by applying a wideband gain value to each short-time segment (306). The compression component also performs predetermined prototype filtering and a QMF filter bank (308) to reduce or eliminate any discontinuities caused by applying different gain values ​​to adjacent segments, as described above. In certain cases, such as the type of audio content or certain characteristics of the audio content, compressing and expanding the audio signal before or after the encoding / decoding stages of an audio codec may degrade rather than enhance the quality of the audio output. In such instances, the compression / expansion process may be turned off or modified to return a different compression / expansion (compression / expansion) level. Thus, the compression component determines (310), among other variables, the appropriateness of the compression / expansion function and / or the optimal level of compression / expansion required for the particular signal input and audio playback environment. This decision step 310 may occur at any practical point in process 300, such as before splitting the audio signal 304 or compressing the audio signal 306. If companding is determined to be appropriate, a gain is applied (306). The encoder then encodes the signal for transmission to the decoder (312) according to the codec's data format. Certain companding control data, such as activation data, synchronization data, companding level data, and other similar control data may be transmitted as part of the bitstream for processing by the enhancement component.

[0036] FIG. 3B is a flowchart illustrating a method for enhancing an audio signal in a post-decoder enhancement component, under one embodiment. As shown in process 350, the decode stage of a codec receives a bitstream encoding the audio signal from the encode stage (352). The decoder then decodes the encoded signal according to the codec data format (353). The enhancement component then processes the bitstream and applies any encoded control data to switch off enhancements or modify enhancement parameters based on the control data (354). The enhancement component divides the audio signal into time segments using an appropriate window shape (356). In one embodiment, the time segments correspond to the same time segments used by the compression component. The enhancement component then calculates appropriate gain values ​​for each segment in the frequency domain and applies the gain values ​​to each time segment to enhance the dynamic range of the audio signal back to the original dynamic range or any other suitable dynamic range.

[0037] Compression / Expansion Control The compression and expansion components, including the compander, of the system 100 are configured to apply pre- and post-processing steps only at certain times during audio signal processing or for certain types of audio content. For example, companders may be beneficial for transient signals in speech and music. However, for other signals, such as static signals, companders may degrade signal quality. Therefore, as shown in FIG. 3A, a compander control mechanism is provided, as shown in block 310, which sends control data from the compression component 104 to the expansion component 114 to regulate the compander operation. In its simplest form, such a control mechanism switches off the compander function for blocks of audio samples for which applying companders would degrade audio quality. In one embodiment, the compander on / off decision is detected in the encoder and transmitted to the decoder as a bitstream element, allowing the compressor and expander to be switched on / off in the same QMF time slot.

[0038] Switching between the two states often introduces a discontinuity in the applied gain, resulting in audible switching artifacts or clicks. Embodiments include mechanisms for reducing or eliminating these artifacts. In a first embodiment, the system can switch the compander function off and on only in frames where the gain is close to 1. In this case, there is only a slight discontinuity between switching and turning the function on and off. In a second embodiment, a third weak compander mode is applied between the on and off modes, i.e., for audio frames between the on and off frames. The weak compander mode slowly transitions the exponential term γ from its default value to 0 during compander operation. As an alternative to the intermediate weak compander mode, the system can implement a start-frame and stop-frame mode. Instead of abruptly switching off the compander function, the compander mode smoothly fades out over a block of audio samples. In a further embodiment, the system is configured to apply an average gain rather than simply switching off the compander. In certain cases, the audio quality of signals without tonal variations can be increased by applying a constant gain factor to an audio frame that is more similar to the gain factor of an adjacent compander-on frame than a constant gain factor of 1.0 in the compander-off state. Such a gain factor can be calculated by averaging all compander-on gains over a frame. Frames containing a constant average compander-on gain are thus signaled in the bitstream.

[0039] It should be noted that although the embodiment is described in the context of a mono audio channel, multiple channels can be easily handled by repeating the application for each channel separately. However, audio signals containing two or more channels present certain additional complexities and are handled by the embodiment of the compander system of Figure 1. The compander strategy should be based on the similarity between the channels.

[0040] For example, it has been observed that in the case of stereo-panned transient signals, independent compression / expansion of individual channels can result in audible image artifacts. In one embodiment, the system determines a single gain value for each time segment from subband samples of both channels and uses the same gain value to compress / expand the two signals. This approach is generally appropriate whenever two channel regions have very similar signals. Here, similarity is determined using, for example, cross-correlation. A detector calculates the similarity between channels and switches between using individual compression / expansion of the channels or joint compression / expansion of the channels. Expanding to more channels involves dividing the channels into groups using similarity criteria and applying joint compression / expansion to the groups. This group information is then transmitted through the bitstream.

[0041] System Implementation FIG. 4 is a block diagram illustrating a system for compressing an audio signal in relation to the encoding stage of a codec, under one embodiment. FIG. 4 shows a hardware circuit or system implementing at least a portion of the compression method for use in the codec-based system shown in FIG. 3A. As shown in system 400, an input audio signal 401 in the time domain is input to a QMF filterbank 402. This filterbank performs an analysis operation that separates the input signal into multiple components, where each bandpass filter conveys a frequency subband of the original signal. Signal reconstruction is performed in a synthesis operation performed by a QMF filterbank 410. In the embodiment of FIG. 4, both the analysis and synthesis filterbanks handle 64 bands. A core encoder 412 receives the audio signal from the synthesis filterbank 410 and generates a bitstream in an appropriate digital format (e.g., MP3, ACC, etc.) by encoding the audio signal.

[0042] The system 400 includes a compressor 406 that applies a gain value to each of the short segments into which the audio signal is divided, producing an audio signal with a compressed dynamic range, such as that shown in FIG. 2B. A compander control unit 404 analyzes the audio signal to determine whether and how much compression should be applied based on the signal type (e.g., speech), signal characteristics (e.g., static vs. transient), or other relevant parameters. The control unit 404 may include a mechanism for detecting temporal peak characteristics of the audio signal. Based on the detected audio signal characteristics and predetermined criteria, the control unit 404 sends an appropriate control signal to the compressor 406 to either turn off the compression function or change the gain value applied to the short segments.

[0043] In addition to compression and expansion, many other coding tools can also operate in the QMF domain. One such tool is advanced apectral extension (A-SPX), shown in block 408 of FIG. 4. A-SPX is a technique used so that perceptually less important frequencies are coded using a coarser coding scheme than more important frequencies. For example, in decoder-side A-SPX, QMF subband samples from lower frequencies are replicated at higher frequencies, and a spectral envelope in the higher frequency bands is then formed using side information transmitted from the encoder to the decoder.

[0044] In a system where both companding and A-SPX are performed in the QMF domain, at the encoder, A-SPX envelope data for higher frequencies may be derived from the uncompressed subband samples, as shown in Figure 4. Compression may then be applied only to the lower frequency QMF samples corresponding to the frequency band of the signal encoded by the core encoder 412. At the decoder 502 of Figure 5, after QMF analysis 504 of the decoded signal, an expansion process 506 is first applied, and an A-SPX operation 508 subsequently regenerates the higher subband samples from the expanded signal at lower frequencies.

[0045] In this embodiment, the QMF synthesis filterbank 410 in the encoder and the QMF analysis filterbank in the decoder 504 together result in a delay of 640-64+1 samples (~9 QMF slots). The core codec delay in this embodiment is 3200 samples (50 QMF slots), for a total delay of 59 slots. This delay is accounted for by embedding control data in the bitstream and using it in the decoder, so that both the compressor in the encoder and the expander in the decoder operate synchronously.

[0046] Alternatively, compression may be applied to the full bandwidth of the original signal at the encoder. An A-SPX envelope may then be derived from the compressed subband samples. In such a case, the decoder first performs A-SPX to reconstruct the full bandwidth of the compressed signal after QMF analysis. An expansion stage is then applied to restore the signal with its original dynamic range.

[0047] Yet another tool that can operate in the QMF domain is the advanced coupling (AC) tool (not shown) in FIG. 4. In an advanced coupling system, two channels are encoded as a mono downmix with additional parametric spatial information that can be applied in the QMF domain at the decoder to reconstruct a stereo output. AC and compression / expansion are used in conjunction with each other. The AC tool can also be placed after the compression stage 406 at the encoder, in which case it would be applied before the expansion stage 506 at the decoder. Alternatively, AC side information can be derived from the uncompressed stereo signal. In that case, the AC tool would operate after the expansion stage 506 at the decoder. A hybrid AC mode is also supported, in which AC is used above a given frequency and discrete stereo is used below this frequency. Or, alternatively, discrete stereo is used above a given frequency and AC is used below this frequency.

[0048] As shown in Figures 3A and 3B, the bitstream transmitted between the encoding and decoding stages of the codec contains certain control data. Such control data constitutes aspect information that allows the system to switch between different compression / expansion modes. Switching control data (for switching compression / expansion on / off) plus potentially several intermediate states may add on the order of one or two bits per channel. Other control data may include signals to determine whether all channels in a discrete stereo or multi-channel configuration use a common compression / expansion gain factor, or whether gain factors should be calculated independently for each channel. Such data may only require one extra bit per channel. Other similar control data elements and their appropriate bit weights may be used according to system requirements and limitations.

[0049] Detection Mechanism In one embodiment, a compander control mechanism is included as part of the compression component 104 to provide control of companders in the QMF domain. The compander control can be configured based on many factors, such as the audio signal type. For example, in most applications, companders should be turned on for speech signals and transient signals, or any other signals in the class of time-peaky signals. The system includes a detection mechanism for detecting signal peaks to aid in generating appropriate control signals for the compander function.

[0050] In one embodiment, for a given core codec, the temporal peak TP(k) over frequency bin k is frame A measurement for is calculated using the following equation:

[0051]

number

[0052] In the above equation, S t where (k) is the subband signal and T is the number of QMF slots corresponding to one core encoder frame. In one embodiment, the value of T may be 32. The temporal peaks calculated for each band can be used to classify sound content into two general categories: static music signals and musical transient or speech signals. TP(k) frame If the value of TP(k) is less than a predetermined value (e.g., 1.2), the signal in that subband of the frame is likely to be a music signal without fluctuations. frame If the value of is greater than this value, the signal is likely to be a musical transient signal or a speech signal. If the value is greater than a higher threshold (e.g., 1.6), the signal is very likely to be a pure musical transient signal, e.g., castanets. Furthermore, it has been observed that for naturally occurring signals, the values ​​of temporal peaks obtained in different bands are more or less similar, and this property can be used to reduce the number of subbands for which temporal peak values ​​need to be calculated. Based on this observation, the system can do one of the following two things:

[0053] In a first embodiment, the detector performs the following process: In a first step, the detector calculates the number of bands with a temporal peak greater than 1.6. In a second step, the detector then calculates the average of the temporal peaks of the bands less than 1.6. If the number of bands found in the first step is greater than 51, or if the average value determined in the second step is greater than 1.45, the signal is determined to be a musical transient signal, and therefore, companders should be switched on. Otherwise, the signal is determined to be one for which companders should not be switched on. Such a detector switches off most of the time for speech signals. In some embodiments, speech signals are often coded by a separate speech coder, and this is generally not an issue. However, in certain cases, it may be desirable to switch on the compander function for speech as well. In this case, the second type of detector would be appropriate.

[0054] In one embodiment, this second type of detector performs the following process: As a first step, the detector calculates the number of bands with a temporal peak greater than 1.2. As a second step, the detector then calculates the average of the temporal peaks of the bands less than 1.2. The detector then applies the following rules: If the result of the first step is greater than 55, turn on companding; if the result of the first step is less than 15, turn off companding; if the result of the first step is between 15 and 55 and the result of the second step is greater than 1.16, turn on companding; if the result of the first step is between 15 and 55 and the result of the second step is less than 1.16, turn off companding. It should be noted that the two types of detectors described are only two examples of many possible solutions for the detection algorithm, and other similar algorithms may also or alternatively be used.

[0055] The companding control function provided by element 404 of FIG. 4 may be implemented in any suitable manner so that companding is or is not used based on a given mode of operation. For example, companding is typically not used on the LFE (low frequency effects) channel of a surround sound system, and is also not used when A-SPX functionality is not implemented (i.e., no QMF). In one embodiment, the companding control function may be provided by a program executed by a circuit or processor-based element, such as companding control element 404. The following is an example of the syntax of a program segment that can implement companding control in one embodiment: Companding_control(nCh) { sync_flag=0; if(nCh>1){ sync_flag } b_needAvg=0 ch_count=sync_flag?1:nCh for(ch=0;ch <ch_count;ch++){ b_compand_on[ch] if(!b_compand_on[ch]){ b_needAvg=1; } } if(b_needAvg){ b_compand_avg: } } The sync_flag, b_compand_on[ch], and b_compand_avg flags or program elements may be on the order of 1 bit long, or may be any other length depending on system limitations and requirements. It should be noted that the program code described above is an example of one way to implement the companding control function, and that other protocols or hardware components may be used to implement the companding control according to some embodiments.

[0056] While the embodiments described so far have included companding programs to reduce quantization noise introduced by the encoder in a codec, it should be noted that aspects of such companding processes may also be applied in a signal processing system that does not include an encoding and decoding (codec) stage. Furthermore, when a companding process is used in conjunction with a codec, the codec may be transform-based or non-transform-based.

[0057] Aspects of the systems described herein can be implemented in a computer-based sound processing network environment suitable for processing digital or digitized audio files. Portions of an adaptive audio system can include one or more networks with any desired number of individual machines. The machines can include one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks can be built on a variety of different network protocols and can be an intranet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0058] One or more components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device associated with the system. It should also be noted that various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media, such as actions, register transfers, logic components, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (fixed), non-volatile storage media in various forms, such as optical, magnetic, or semiconductor storage media.

[0059] Unless the context clearly requires otherwise, throughout the specification and claims, the terms "comprise," "comprising," and the like, are to be understood in an inclusive sense as opposed to an exclusive or exhaustive sense. Terms using one or more numbers also include the plural or single number, respectively. In addition, the terms "herein," "hereunder," "above," "below," and words of similar import refer to this application as a whole and not to any particular portions of this application. When the term "or" is used in connection with a list of two or more items, it is intended to cover all of the following interpretations of the term: every item in the list, all items in the list, and every combination of items in the list.

[0060] While one or more embodiments have been described by way of example and with reference to specific embodiments, it is to be understood that the one or more implementations are not limited to the disclosed embodiment. On the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements. The following notes are provided regarding the above embodiment. (Appendix 1) 1. A method for enhancing an audio signal, comprising: receiving an audio signal; and expanding the audio signal to an extended dynamic range by an expansion process, The expansion process comprises: dividing the received audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain for each time segment in the frequency domain using a non-energy based average of a frequency domain representation of the audio signal; applying a separate gain value to each time segment to obtain the extended dynamic range; applying the individual gain values ​​amplifies relatively high intensity segments and attenuates relatively low intensity segments; method. (Appendix 2) The segments are overlapping. The method described in Appendix 1. (Appendix 3) a first filterbank is used to analyze the audio signal to obtain a frequency domain representation; and the determined window shape corresponds to a prototype filter for the first filter bank; The method described in Appendix 2. (Appendix 4) the first filter bank is one of a quadrature modulated filter (QMF) bank or a short-time Fourier transform. The method described in Appendix 3. (Appendix 5) the wideband gain for each time segment is calculated using the subband samples in a subset of subbands in the each time segment. The method described in Appendix 3. (Appendix 6) the subset of subbands corresponds to all frequency bands spanned by the first filterbank. The method described in Appendix 5. (Appendix 7) the gain for each time segment is derived from the p-norm of the subband samples in each time segment; where p is a positive real number not equal to 2. The method described in Appendix 5. (Appendix 8) the wideband gain is applied in the region of the first filter bank; The method described in Appendix 5. (Appendix 9) each wideband gain value is calculated from a first subset of subbands of the first filter bank and applied to a second subset of subbands of the first filter bank; wherein the second set of subbands comprises the first subset of subbands. The method described in Appendix 8. (Appendix 10) the first and second subsets of subbands are identical and correspond to a low frequency region of the audio signal. The method described in Appendix 9. (Appendix 11) the first subset of subbands corresponds to a low frequency region of the audio signal; and the second subset of subbands corresponds to all frequency bands spanned by the first filterbank. The method described in Appendix 9. (Appendix 12) the received audio signal has previously been compressed by a process; The process comprises: receiving an initial audio signal; compressing said first audio signal by a compression process to substantially reduce its original dynamic range; The compression process comprises: dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain for each segment using a non-energy based average of frequency domain samples of the initial audio signal; applying a gain value calculated from the original audio signal to each segment of the plurality of segments to amplify segments of relatively low intensity and attenuate segments of relatively high intensity; 4. The method of claim 3, comprising: (Appendix 13) the wideband gain calculated by the expansion process is substantially the inverse of the wideband gain calculated by the compression process for the corresponding time segment; 12. The method described in Appendix 12. (Appendix 14) The wideband gain is calculated in the compression process to analyze the original audio signal to obtain a frequency domain representation; and the defined window shape for division is the same as a prototype filter for the first filter bank; and the second filter bank is identical to the first filter bank; 12. The method described in Appendix 12. (Appendix 15) the received signal for the enhancement process is obtained after modification of the compressed signal by an audio encoder generating a bitstream and a decoder decoding the bitstream; 12. The method described in Appendix 12. (Appendix 16) the audio encoder and the decoder are both transform-based; and the time segments of the audio signal in the compression and expansion processes are substantially shorter than one window length of a transform in the audio encoder and decoder; The method described in Appendix 15. (Appendix 17) The method further comprises: generating control information that determines the operational state of the expansion process; transmitting said control information in a bitstream transmitted from said encoder to said decoder; 16. The method of claim 15, comprising: (Appendix 18) dividing the audio signal in the bitstream into frames with each frame according to a plurality of time segments of the enhancement process; the operating state is selected from a group; The group is applying the expansion process to each time segment in a frame; The expansion process does not apply to every time segment in a frame. applying the expansion process to each time segment in a frame using a modified gain calculation, wherein the gain applied at each time segment is the average gain of all time segments in the frame; applying the expansion process to each time segment in a frame using a modified gain calculation, the calculation resulting in a gain value intermediate to when not applying the expansion process at all; using a stop frame to fade out frames to which the enhancement process is applied and fade in frames to which the enhancement process is not applied; using a start frame to fade out from frames to which the enhancement process is not applied and fade in to frames to which the enhancement process is applied; and applying the expansion process completely; 18. The method of claim 17, comprising: (Appendix 19) the control information for the enhancement process is determined by the compression step based on one or more characteristics of the original audio signal, including at least one of a content type of the audio signal and stationary versus transient characteristics of the audio signal; 18. The method described in Appendix 18. (Appendix 20) the control information is determined such that switching between operating states minimizes the occurrence of signal discontinuities; 19. The method described in Appendix 19. (Appendix 21) The control information also controls the compression process, and having the effect of turning off the compression process when the expansion process is switched off and turning on the compression process when the expansion process is switched on, allowing for a corrected gain calculation for the expansion if a corrected gain calculation for the expansion is made; Use a stop frame when a stop frame is used in the expander, and use a start frame when a start frame is used in the expander; 21. The method described in Appendix 20. (Appendix 22) the compressed audio signal and the audio signal received by the expander have a number, N, channels, where N is greater than 1; the channels are grouped into one or more disjoint subsets; the groupings at the compressor and the expander are identical; the channels in each group are compressed by sharing the same gain in the compressor and expanded by sharing the same gain in the expander; The method described in Appendix 15. (Appendix 23) the grouping is predefined and known to the compressor and the expander; 23. The method described in Appendix 22. (Appendix 24) Each group contains exactly one channel, and there are N groups. 24. The method described in Appendix 23. (Appendix 25) The grouping of channels may include: calculating a similarity metric between channels in the compressor; grouping similar channels together based on said similarity metric; transmitting information of the grouping through the bitstream; 23. The method of claim 22, comprising: (Appendix 26) encoding at least two channels as a mono downmix with additional parametric spatial information applied in the first filterbank domain to reconstruct a stereo output; the additional parametric spatial information is either used on a predetermined frequency with separate stereo information used under the predetermined frequency, or used under a predetermined frequency with separate stereo information used on the predetermined frequency, 23. The method described in Appendix 22. (Appendix 27) 1. A method for compressing an audio signal, comprising: receiving an initial audio signal; and substantially reducing the dynamic range of the initial audio signal by a compression process; The compression process comprises: dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy based average of frequency domain samples of the first audio signal; applying a separate gain value to each segment of the plurality of segments to amplify relatively low intensity segments and attenuate relatively high intensity segments; method. (Appendix 28) the segments are overlapping; a first filterbank is used to analyze the audio signal to obtain a frequency domain representation; and the determined window shape corresponds to a prototype filter for the first filter bank; The method described in Appendix 27. (Appendix 29) the first filter bank is one of a quadrature modulated filter (QMF) bank or a short-time Fourier transform. 29. The method described in Appendix 28. (Appendix 30) Each individual gain value is calculated using subband samples in a subset of subbands in each time segment. 29. The method described in Appendix 28. (Appendix 31) the subset of subbands corresponds to all frequency bands spanned by the first filterbank; and the gain is applied in the domain of the first filter bank; 31. The method described in Appendix 30. (Appendix 32) the gain for each time segment is derived from the p-norm of the subband samples in each time segment; where p is a positive real number not equal to 2. 31. The method described in Appendix 30. (Appendix 33) the gains are calculated from a first subset of subbands of the first filter bank and applied to a second subset of subbands of the first filter bank; wherein the second set of subbands comprises the first subset of subbands. 31. The method described in Appendix 30. (Appendix 34) the first and second subsets of subbands are identical and correspond to a low frequency region of the audio signal. 34. The method described in Appendix 33. (Appendix 35) the first subset of subbands corresponds to a low frequency region of the audio signal; and the second subset of subbands corresponds to all frequency bands spanned by the first filterbank. 34. The method described in Appendix 33. (Appendix 36) The method further comprises: sending a compressed version of the initial audio signal to an expansion component that performs an expansion process; The expansion process comprises: receiving the compressed version of an audio signal; expanding the compressed version of the audio signal by a process to substantially restore the original dynamic range of the audio signal; The process comprises: dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy based average of a frequency domain representation of the first audio signal; applying a separate gain value of the wideband gain to each time segment so as to amplify segments of relatively high intensity and attenuate segments of relatively low intensity; Including, The method described in Appendix 27. (Appendix 37) the gain calculated by the compression step is substantially the inverse of the gain calculated by the expansion process for the same time segment. The method described in Appendix 36. (Appendix 38) a second filter bank is used in the extension process to analyze the initial audio signal to obtain a frequency domain representation; and the defined window shape for the division is identical to the prototype filter for the filter bank, and the second filter bank is identical to the first filter bank; The method described in Appendix 36. (Appendix 39) the received signal for the expansion step is obtained after modification of the compressed signal by an audio encoder generating a bitstream and a decoder decoding the bitstream; The method described in Appendix 36. (Appendix 40) the audio encoder and the decoder are both transform-based; and the time segments of the audio signal in the compressing and expanding steps are substantially shorter than one window length involved in a transform in the audio encoder and decoder; 39. The method described in Appendix 39. (Appendix 41) The method further comprises: generating control information that determines an operating state of the expansion step; transmitting said control information in a bitstream transmitted from said encoder to said decoder; 39. The method of claim 39, comprising: (Appendix 42) dividing the audio signal in the bitstream into frames with each frame according to a plurality of time segments of the enhancement process; the operating state is selected from a group; The group is applying the expansion process to each time segment in a frame; The expansion process does not apply to every time segment in a frame. applying the expansion process to each time segment in a frame using a modified gain calculation, wherein the gain applied at each time segment is the average gain of all time segments in the frame; applying the expansion process to each time segment in a frame using a modified gain calculation, the calculation resulting in a gain value intermediate to when not applying the expansion process at all; using a stop frame to fade out frames to which the enhancement process is applied and fade in frames to which the enhancement process is not applied; using a start frame to fade out from frames to which the enhancement process is not applied and fade in to frames to which the enhancement process is applied; and applying the expansion process completely; 42. The method of claim 41, comprising: (Appendix 43) the control information for the enhancement process is determined by the compression step based on one or more characteristics of the original audio signal, including at least one of a content type of the audio signal and stationary versus transient characteristics of the audio signal; 42. The method described in Appendix 42. (Appendix 44) the control information is determined such that switching between operating states minimizes the occurrence of signal discontinuities; 43. The method described in Appendix 43. (Appendix 45) The control information also controls the compression process, and having the effect of turning off the compression process when the expansion process is switched off and turning on the compression process when the expansion process is switched on, allowing for a corrected gain calculation for the expansion if a corrected gain calculation for the expansion is made; Use a stop frame when a stop frame is used in the expander, and use a start frame when a start frame is used in the expander; The method described in Appendix 44. (Appendix 46) the compressed audio signal and the audio signal received by the expander have a number, N, channels, where N is greater than 1; the channels are grouped into one or more disjoint subsets; the groupings at the compressor and the expander are identical; the channels in each group are compressed by sharing the same gain in the compressor and expanded by sharing the same gain in the expander; 39. The method described in Appendix 39. (Appendix 47) the grouping is predefined and known to the compressor and the expander; 46. ​​The method described in Appendix 46. (Appendix 48) Each group contains exactly one channel, and there are N groups. The method described in Appendix 47. (Appendix 49) The grouping of channels may include: calculating a similarity metric between channels in the compressor; grouping similar channels together based on said similarity metric; transmitting information of the grouping through the bitstream; 47. The method of claim 46, comprising: (Appendix 50) encoding at least two channels as a mono downmix with additional parametric spatial information applied in the first filterbank domain to reconstruct a stereo output; the additional parametric spatial information is either used on a predetermined frequency with separate stereo information used under the predetermined frequency, or used under a predetermined frequency with separate stereo information used on the predetermined frequency, 49. The method described in Appendix 49. (Appendix 51) 1. An apparatus for compressing an audio signal, comprising: a first interface for receiving a first audio signal; a compressor for compressing the initial audio signal to substantially reduce an original dynamic range of the initial audio signal; The compressor comprises: Dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy-based average of frequency domain samples of the first audio signal; applying a separate gain value to each of the plurality of segments to amplify relatively low intensity segments and attenuate relatively high intensity segments; A device that performs compression by (Appendix 52) The apparatus further comprises: a first filterbank for analyzing the audio signal to obtain a frequency domain representation; the determined window shape corresponds to a prototype filter for the first filter bank; and the first filter bank is one of a quadrature modulated filter (QMF) bank or a short-time Fourier transform. 52. The apparatus of claim 51. (Appendix 53) The individual gain values ​​are calculated using subband samples in the subset of subbands in each time segment. 53. The apparatus of claim 52. (Appendix 54) the subset of subbands corresponds to all frequency bands spanned by the first filterbank; and the gain is applied in the domain of the first filter bank; 54. The apparatus of claim 53. (Appendix 55) The apparatus further comprises: a second interface for transmitting a compressed version of the first audio signal to an expander; The expander comprises: receiving said compressed version of the audio signal; to restore the compressed version of the audio signal to substantially the original dynamic range of the audio signal; Dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy-based average of a frequency domain representation of the first audio signal; applying a separate gain value of the wideband gain to each time segment to amplify segments of relatively high intensity and attenuate segments of relatively low intensity; Expand by 53. The apparatus of claim 52. (Appendix 56) the gain calculated by the compressor is substantially the inverse of the gain calculated by the expander for the same time segment. 56. The apparatus of claim 55. (Appendix 57) The apparatus further comprises: a second filter bank that analyzes the first audio signal to obtain a frequency domain representation thereof; the defined window shape for the division is identical to the prototype filter for the filter bank, and the second filter bank is identical to the first filter bank; 56. The apparatus of claim 55. (Appendix 58) The apparatus further comprises: an audio codec encoding stage and a decoding stage configured to transmit a compressed version of the audio signal from a compressor to an expander; the encoder and decoder are both transform-based; 56. The apparatus of claim 55. (Appendix 59) The apparatus further comprises: a control component that generates control information that determines an operational state of the extender and transmits the control information in the bitstream; the control information for the enhancement process is determined by the compression step based on one or more characteristics of the original audio signal, including at least one of a content type of the audio signal and stationary versus transient characteristics of the audio signal; 59. The apparatus of claim 58. (Appendix 60) The apparatus further comprises: a parametric spatial information component that applies the parametric spatial information in the first filterbank domain to reconstruct a stereo output; the parametric spatial information is either used on a predetermined frequency with separate stereo information used under the predetermined frequency, or used under a predetermined frequency with separate stereo information used over the predetermined frequency, 56. The apparatus of claim 55. (Appendix 61) 1. An apparatus for enhancing an audio signal, comprising: a first interface for receiving a compressed audio signal; an expander for substantially restoring the compressed audio signal to its original uncompressed dynamic range; The expander comprises: Dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy-based average of frequency domain samples of the first audio signal; applying a separate gain value to each segment of the plurality of segments to amplify relatively high intensity segments and attenuate relatively low intensity segments; A device that performs expansion by (Appendix 62) The apparatus further comprises: a first filterbank for analyzing the audio signal to obtain a frequency domain representation; the determined window shape corresponds to a prototype filter for the first filter bank; and the first filter bank is one of a quadrature modulated filter (QMF) bank or a short-time Fourier transform. 62. The apparatus of claim 61. (Appendix 63) the wideband gain includes an individual gain value for each time segment; and Each individual gain value is calculated using subband samples in a subset of subbands in each time segment. 63. The apparatus of claim 62. (Appendix 64) the subset of subbands corresponds to all frequency bands spanned by the first filterbank; and the gain is applied in the domain of the first filter bank; 64. The apparatus of claim 63. (Appendix 65) The apparatus further comprises: a second interface for receiving the compressed audio signal from a compressor that receives the original audio signal; The compressor comprises: to substantially reduce the original dynamic range of the first audio signal; Dividing the initial audio signal into a plurality of time segments using a defined window shape; calculating a wideband gain in the frequency domain using a non-energy-based average of frequency domain samples of the first audio signal; applying a respective gain value to each time segment of the plurality of segments to amplify relatively low intensity segments and attenuate relatively high intensity segments; compressing the first audio signal by 63. The apparatus of claim 62. (Appendix 66) the gain calculated by the compressor is substantially the inverse of the gain calculated by the expander for the same time segment. 66. The apparatus of claim 65. (Appendix 67) The apparatus further comprises: a second filter bank that analyzes the first audio signal to obtain a frequency domain representation thereof; the defined window shape for the division is identical to the prototype filter for the filter bank, and the second filter bank is identical to the first filter bank; 66. The apparatus of claim 65. (Appendix 68) The apparatus further comprises: an audio codec encoding stage and a decoding stage configured to transmit a bitstream of a compressed version of the audio signal from a compressor to an expander; the encoder and decoder are both transform-based; 66. The apparatus of claim 65. (Appendix 69) The apparatus further comprises: a control component that generates control information that determines an operational state of the extender and transmits the control information in the bitstream; the control information for the enhancement process is determined by the compression step based on one or more characteristics of the original audio signal, including at least one of a content type of the audio signal and stationary versus transient characteristics of the audio signal; 69. The apparatus of claim 68. (Appendix 70) The apparatus further comprises: a parametric spatial information component that applies the parametric spatial information in the first filterbank domain to reconstruct a stereo output; the parametric spatial information is either used on a predetermined frequency with separate stereo information used under the predetermined frequency, or used under a predetermined frequency with separate stereo information used over the predetermined frequency, 66. The apparatus of claim 65. [Explanation of symbols]

[0061] 104 Compression Components 106 Encoder 110 Network 112 decoder 114 Extended Components 116 Audio Output 406 Compressor 412 Core Encoder

Claims

1. 1. A method for compressing an audio signal containing multiple channels, comprising: receiving, by a computer, a time-frequency tiled representation of an audio signal; The time-frequency tiled representation of the audio signal divides the audio signal into time slots, and the time-frequency tiled representation of the audio signal is a representation in the QMF domain, each time slot being divided into frequency sub-bands, the frequency sub-bands being uniformly spaced apart; and compressing, by the computer, the time-frequency tiled representation of the audio signal; Reducing the dynamic range of the audio signal; Including, The step of compressing the time-frequency tiled representation of the audio signal comprises the steps of: dividing the channels of the audio signal into distinct subsets of channels based on grouping information; For each distinct subset of channels, calculating a sharing gain for a time slot of the time-frequency tiled representation of the audio signal, wherein calculating the sharing gain comprises reducing a compression level in response to control data; applying a shared gain for the time slot to each frequency subband of each channel of the respective subset of channels; A method comprising:

2. A non-transitory computer-readable storage medium containing instructions, The instructions, when executed by one or more processors, perform the method of claim 1. A non-transitory computer-readable storage medium.

3. 1. An apparatus for compressing an audio signal comprising multiple channels, comprising: a first interface for receiving a time-frequency tiled representation of an audio signal; a first interface, wherein the time-frequency tiled representation of the audio signal divides the audio signal into time slots, and the time-frequency tiled representation of the audio signal is a representation in the QMF domain, each time slot being divided into frequency sub-bands, the frequency sub-bands being uniformly spaced; a compressor for compressing time-frequency tiled representations of the audio signal, a compressor for reducing the dynamic range of the audio signal; Including, Compressing the time-frequency tiled representations of the audio signal comprises: dividing the channels of the audio signal into distinct subsets of channels based on grouping information; For each distinct subset of channels, calculating a sharing gain for a time slot of the time-frequency tile representation of the audio signal, wherein calculating the sharing gain comprises reducing a compression level in response to control data; applying a shared gain for the time slot to each frequency subband of each channel of the respective subset of channels; 1. An apparatus comprising: