Dynamic range compression with reduced artifacts

The dynamic range compression method addresses artifacts in audio signals by applying reduced DRC based on signal averaging and controlled release times, optimizing loudness and reducing distortion in systems with limited power processing.

JP7771343B2Active Publication Date: 2025-11-17DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024228117
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2024-12-25
Publication Date
2025-11-17
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

Existing dynamic range compression (DRC) systems introduce undesirable artifacts such as 'pumping' and 'breathing' when applied to audio signals, particularly in systems with limited power processing capabilities, and fail to optimize loudness and prevent distortion.

Method used

Implementing a dynamic range compression method that applies reduced DRC or no DRC when the average loudness of an audio signal approaches a target level, controlling release time constants based on transient presence, and using slow-smoothed averages to determine gain adjustments, thereby reducing artifacts.

Benefits of technology

Reduces or prevents pumping and breathing artifacts while maintaining optimal loudness and preventing distortion in systems with limited power processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771343000001
    Figure 0007771343000001
  • Figure 0007771343000002
    Figure 0007771343000002
  • Figure 0007771343000003
    Figure 0007771343000003
Patent Text Reader

Abstract

To provide a method and a system for performing dynamic range compression (DRC) on audio in a manner intended to reduce or prevent undesirable artifacts in the output audio, such as pumping and breathing.SOLUTION: A DRC system performs DRC to maximize the average loudness and reduce or prevent distortion during playback. When the average loudness of the input audio approaches, meets, or exceeds a target, knee point for DRC or a signal level near the maximum playback level of the intended playback system, then it is assumed that such input audio is already compressed and therefore applies reduced DRC or full DRC to the input audio.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to European Patent Application No. 19197154.8, filed September 13, 2019, and U.S. Provisional Patent Application No. 62 / 899,769, filed September 13, 2019, which are incorporated herein by reference.

[0002] Technical Field The present disclosure relates to dynamic range compression of audio signals. [Background technology]

[0003] Herein, dynamic range compression is sometimes referred to as "DRC" and dynamic range compressors are sometimes referred to as "DRCs."

[0004] 1, a traditional dynamic range compressor (“DRC”) includes a level estimator 1, a dynamic range compression (DRC) gain determination subsystem 6, and a gain application subsystem 7 coupled as shown. In some embodiments, gain determination subsystem 6 includes a smoother 3 and a dynamic range compression (DRC) gain curve subsystem 5 coupled as shown. Subsystem 7 is coupled and configured to apply a time-varying gain g(t) (e.g., a series of gain values) output from subsystem 6 (e.g., from subsystem 5 of subsystem 6 as shown in FIG. 1) to an input audio signal to generate an audio output signal.

[0005] The dynamic range compression applied to the input audio signal may decrease the power of segments of the input audio signal whose smoothed level (output from smoother 3) is above a threshold (first knee point) (i.e., subsystem 7 applies segment gain(s) determined by subsystem 6 below gain(s) 1) and increase the power of segments of the input audio signal whose smoothed level (output from smoother 3) is below a second threshold (the "lower" knee point) (i.e., subsystem 7 may apply a gain greater than 1 determined by subsystem 6 to such segments).

[0006] Dynamic range compression (DRC) has an attack period (starting at or near the point where the input audio level (e.g., as indicated by the output of Level Estimator 1) rises to the knee point, with a duration known as the attack time) and a release period (starting at or near the point where the input audio level falls to the knee point, with a duration known as the release time). Full DRC is applied during the period after the attack but before the release. The amount of DRC applied can increase from zero (when Subsystem 6 outputs Gain 1) to the full amount during the attack, then return to zero during the release.

[0007] One important possibility with DRC is that there is a large gain for low-level input audio below the knee point, and then as the input signal level approaches its maximum level, the gain decreases monotonically to a gain of 1. Thus, DRC does not actually reduce the higher levels of the input audio, but instead only increases the lower levels of the input audio. Such cases arise when DRC is used to obtain maximum loudness from a low-power system.

[0008] When the input audio signal jumps to a maximum or minimum level, this change is indicated by the output of level estimator 1 in FIG. 1, but the DRC gain (output from subsystem 6) typically does not immediately change to its full value (i.e., the value it would have if the input audio signal level had remained fixed at the maximum or minimum level and had not jumped to that maximum or minimum level). Instead, the DRC gain changes smoothly to its full value (e.g., due to the presence of smoother 3 in FIG. 1). The time required for the DRC gain to reach its full value (after the input signal level jumps to the maximum level) is the attack time, and the time required for the DRC gain to reach its full value (after the input signal level jumps to the minimum level) is the release time. Because the DRC gain typically approaches its full value exponentially, when attack and release times are quoted (e.g., in system specifications), the quoted attack (or release) time is often the time required for the DRC gain to reach partway toward its full value.

[0009] In the implementation shown in FIG. 1, smoother 3 (i.e., the smoothing time constant used by smoother 3) determines the attack and release times (the attack time may be equal to or different from the release time, or one or both may be zero).

[0010] In other implementations of subsystem 6 (not specifically shown in FIG. 1 ), smoother 3 is replaced by another element (e.g., a smoother operating on the gain values ​​output from DRC gain curve subsystem 5) or a subsystem that determines the attack and release times for each interval in which DRC is applied by subsystem 6. Generally, subsystem 6 is configured to determine (e.g., in response to user selection) the attack and release times for each interval in which subsystem 6 applies DRC (i.e., each interval in which subsystem 6 outputs a gain that is not unity).

[0011] In some implementations of dynamic range compression, the compression performance is split, with a gain application subsystem (e.g., subsystem 7) implemented in the decoder or playback system or device, and other elements of the compression (e.g., subsystems 1 and 6) implemented in the encoder, with the gain g(t) being sent as metadata (to the decoder or playback system or device) along with the input audio in the encoded bitstream. Some embodiments of the present invention (described below) contemplate such implementations.

[0012] Level estimator 1 is coupled and configured to determine and provide a level estimate to subsystem 6 (e.g., to smoother 3 in the implementation shown in FIG. 1 ). The level estimate is an estimate (typically varying over time) of the loudness of the input audio signal (e.g., the level estimate represents a sequence of average level or average power values, each averaging time long enough for stability of the dynamic range compression applied by the system of FIG. 1 ). One typical level estimate is average power. Another example of a level estimate is loudness as defined by the ITU-R BS.1770 loudness standard. Smoothener 3 is coupled and configured to apply smoothing to the level estimate output from estimator 1 and generate (and assert to subsystem 5) a smoothed estimate of the average level or power of the input audio signal (the smoothed level estimate).

[0013] In the implementation shown in FIG. 1 , in response to the smoothed level estimates (e.g., a smoothed sequence of mean level or power values) determined by smoother 3, subsystem 5 determines a sequence of values ​​of gain g(t). Subsystem 5 implements a function (typically referred to as a “DRC gain curve”) that maps each value (of mean level or power) output from smoother 3 to a value (gain value) of gain g(t). Gain element 7 applies gain g(t) to the input audio signal to generate an output audio signal (which is a dynamic range-compressed version of the input audio signal), e.g., by applying each value (of the sequence of values) of gain g to a corresponding value of the input audio signal (e.g., its sequence of values).

[0014] In some other implementations, subsystem 6 applies a DRC gain curve that maps each value (of average level or power) output from level estimator 1 (rather than from smoother 3 as shown in FIG. 1 ) to a value of gain g(t) (gain value) (including by implementing DRC attack and release (with determined attack and release times for each interval of DRC application by subsystem 6)), and the gain value g(t) (optionally modified to implement the attack or release) is provided to gain element 7. Summary of the Invention [Means for solving the problem]

[0015] Some embodiments of the present invention are methods for generating output audio for (e.g., optimized for) playback by a system or device (e.g., a notebook, laptop, tablet, sound bar, mobile phone, or other device including or for use with small speakers) having limited power processing capabilities, preferably performing dynamic range compression (DRC) on an audio signal in a manner intended to also reduce or prevent the occurrence of undesirable artifacts in the output audio (e.g., known as "pumping" and "breathing"). In some embodiments, the DRC is performed (or is intended to) maximize average loudness (or provide a sufficiently large average loudness) during playback (or while preventing the loss of quieter elements of the audio) and also reduce or prevent distortion (e.g., to reduce or prevent the occurrence of pumping and / or breathing artifacts and / or to reduce or prevent timbre changes due to frequency components produced by the nonlinear gain application of the DRC). Some embodiments perform DRC on audio signals in a manner intended to optimize content for radio broadcast or the general audibility of audio content or components within an audio stream.

[0016] Herein, the expression "DRC application time" may be used to refer to the attack time (or release time) of an instance of application of DRC (e.g., an instance of application that applies a gain that is not unity, or an instance of application in which the DRC gain curve determines a gain that is not unity after the attack and before the release), or the duration of such an instance of application of DRC (including the attack and release).

[0017] In a first class of embodiments of the dynamic range compression (DRC) method of the present invention, a reduced DRC gain (e.g., no DRC) is applied to an input audio signal when its average loudness (e.g., average level or power) approaches (or matches or exceeds) a target. This is because such an input audio signal is assumed to be already compressed (e.g., maximizing loudness while preventing the loss of quieter elements of the audio during playback). Otherwise, a full DRC is applied to the input audio signal. Applying a reduced DRC (e.g., no DRC) when a full DRC is unnecessary reduces or prevents the occurrence of pumping and / or breathing artifacts that would result from a full DRC. The average loudness is determined over a longer (e.g., much longer) time period (an "averaging" time) than the DRC application time. The target may be a target signal level or a target signal power. In typical embodiments of the first class, the target is an audio signal level that is close to (e.g., equal to or substantially equal to) the knee point for the DRC or the maximum playback level of the playback system or device that reproduces the output audio.

[0018] A second class of embodiments of the dynamic range compression (DRC) method of the present invention is directed to reducing pumping artifacts during DRC for input audio having regular transients (e.g., a sequence of identical or similar transients). Exemplary embodiments of the second class control the release time constant (of each application of dynamic range compression). This includes implementing a first release time constant (referred to as a relatively slow release time constant) when a segment of the input audio signal contains regular transients (including by applying a smoothed dynamic range compression gain to that segment of the input audio signal) and implementing a relatively fast release time constant (i.e., a release time constant faster than the first release time constant) when a different segment of the input audio signal does not contain regular transients (including by applying an unsmoothed dynamic range compression gain to that different segment of the input audio signal). When the relatively slow release time constant is implemented, pumping artifacts are reduced or prevented from occurring.

[0019] A third class of dynamic range compression (DRC) method embodiments of the present invention are directed to reducing breathing artifacts during DRC for decaying input audio. Exemplary embodiments of the third class control the release time constant (of each application of dynamic range compression) in response to the loudness gradient of the input audio signal. This control typically implements a faster release time constant in response to increased steepness of the loudness gradient (to reduce or prevent the occurrence of breathing artifacts) and a slower release time constant in response to decreased steepness of the loudness gradient (to reduce or prevent the occurrence of pumping artifacts).

[0020] Another aspect of the invention is a system (e.g., a dynamic range compressor) or device configured to perform any embodiment of the method of the invention on an input audio signal. In one class of embodiments, the invention is a playback system (e.g., a notebook, laptop, tablet, sound bar, mobile phone, or other device having (or for use with) small speakers, or having limited (e.g., physically limited) power processing capability) configured to perform dynamic range compression (in accordance with any embodiment of the method of the invention) to generate dynamically range compressed audio, and to perform playback of the dynamically range compressed audio.

[0021] In some embodiments, a system of the present invention is or includes a general-purpose or special-purpose processor that is programmed with software (or firmware) and / or otherwise configured to perform method embodiments of the present invention. In some embodiments, a system of the present invention is a general-purpose processor that is coupled to receive input audio data and programmed (with appropriate software) to perform method embodiments of the present invention to generate output audio data. In some embodiments, a system of the present invention is a digital signal processor that is coupled to receive input audio data and configured (e.g., programmed) to perform method embodiments of the present invention to generate output audio data in response to the input audio data.

[0022] Aspects of the present invention include a system configured (e.g., programmed) to perform any of the method embodiments of the present invention, and a computer-readable medium (e.g., disk) storing code for performing any of the method embodiments of the present invention. [Brief explanation of the drawings]

[0023] [Figure 1]1 is a block diagram of a conventional system configured to perform dynamic range compression on an input audio signal.

[0024] [Figure 2] FIG. 1 is a block diagram of an embodiment of the dynamic range compression system of the present invention.

[0025] [Figure 2A] FIG. 2 is a block diagram of another embodiment of the dynamic range compression system of the present invention.

[0026] [Figure 3] FIG. 2 is a block diagram of another embodiment of the dynamic range compression system of the present invention.

[0027] [Figure 4] FIG. 2 is a block diagram of another embodiment of the dynamic range compression system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0028] Some embodiments of the present invention provide improvements and technical solutions for reducing or preventing the occurrence of undesirable (e.g., annoying) artifacts known as "pumping" and "breathing" as a result of dynamic range compression. Different classes of embodiments implement different approaches (described herein) for preventing or reducing such pumping and breathing artifacts.

[0029] A first class of embodiments of the dynamic range compressor and dynamic range compression (DRC) method of the present invention will be described with reference to FIG. 2. In some embodiments of this class, the DRC applies (or the dynamic range compressor is configured to apply) a reduced DRC (or no DRC) to the input audio signal when the average level (or power) of the input audio signal approaches (or matches or exceeds) a target (e.g., a high average level), since such an input audio signal is assumed to be already compressed. The average level (or power) of the input audio signal is determined over a time ("averaging time") that is longer (e.g., much longer) than each attack time and / or each release time of the dynamic range compression, and the reduced DRC is applied "when" the signal has the average level (or power). This means that the reduced DRC is applied to each segment (of a duration equal to or greater than the averaging time) of the input audio signal that has the average level (or power). With reference to the first class of embodiments, the term "target" is used broadly to refer to a target signal level (or signal power) such that it can be reasonably assumed that the input audio signal has already been compressed so that its average level (or power) approaches (or meets or exceeds) the target. In typical embodiments of the first class, the target level is an audio signal level that is near the knee point for DRC or the maximum playback level of the playback system or device that reproduces the output audio.

[0030] The dynamic range compressor of Figure 2 is an exemplary embodiment of a first class of embodiments. The compressor of Figure 2 differs from the conventional dynamic range compressor of Figure 1 in that the compressor of Figure 2 includes a slow smoother 2 and a gain adjustment subsystem 4 coupled as shown (while the compressor of Figure 1 does not). The other elements of the DRC of Figure 2 (elements 1, 6, and 7) are identical to the corresponding (identically numbered) elements of the DRC of Figure 1. Subsystem 6 of Figure 2 may be implemented in any of the ways that subsystem 6 of Figure 1 may be implemented (e.g., including elements 3 and 5 of Figure 1).

[0031] A typical DRC gain curve (e.g., implemented by subsystem 5 of FIG. 1 or FIG. 2) determines the following gain values ​​(values ​​of gain g(t) in FIG. 1 or FIG. 2): In response to the audio signal having a high average level or power (e.g., the smoothed level output from the smoother 3) being above a threshold (knee point), 1 in response to the audio signal having an average level or power (e.g., a smoothed level output from the smoother 3) below the knee point but above a second threshold (second knee point), and a value greater than 1 in response to the audio signal having a low average level or power below the second threshold; As a result, the DRC (which generates the output audio signal in response to the input audio signal) maintains a monotonic increase in the average output audio signal level (or power) as the average input signal level (or power) increases. The gain value determined by the DRC gain curve (and output from subsystem 6 in Figure 1 or Figure 2) dynamically changes as the average level (or power) of the input audio signal changes. However, this can result in undesirable artifacts (e.g., pumping artifacts) in the output audio signal, especially when the average level (or power) of the input signal is near the knee point of the DRC gain curve.

[0032] The inventors have come to realize that if an input audio signal is already sufficiently compressed (and can be reproduced with sufficient loudness by the intended reproduction device), an ideal DRC system would apply less DRC (e.g., no DRC) to such input audio signals than it applies to other input audio signals. In other words, an ideal DRC system would be "unobtrusive" (and therefore not introduce significant pumping artifacts) to sufficiently compressed (and sufficiently loud) input audio signals. The inventors have further come to realize that if an input audio signal has an average level or power that approaches (or matches or exceeds) an appropriate target value (e.g., a target value that is high enough that it can be reasonably assumed that the input audio signal has already been compressed and can be reproduced with sufficient loudness by the intended reproduction device), and if the average is determined over a sufficiently long time interval (i.e., a period that is longer (e.g., much longer) than each attack time and / or each release time of the dynamic range compression), then the DRC system should apply less DRC to such input audio signal than it applies to other input audio signals (i.e., the DRC system should "stay out of the way" and therefore not introduce significant pumping artifacts into the input audio signal).

[0033] Referring again to Figure 2, slow smoother 2 generates (and provides to gain adjustment subsystem 4) a slowly smoothed version of the level (or power) estimate output by level estimator 1. The output of slow smoother 2 is an estimate of the average level or power of the input audio signal, the average being over a longer (e.g., much longer) time than the attack (and / or release) time employed by subsystem 6 (or than the duration of each application, including attack and release, of a non-unity gain DRC by subsystems 6 and 7).

[0034] Herein, the expression "DRC application time" may be used to refer to the attack time (or release time) of an instance of application of DRC (e.g., an instance of application that applies a gain that is not unity, or an instance of application in which the DRC gain curve determines a gain that is not unity after the attack and before the release), or the duration of such an instance of application of DRC (including the attack and release).

[0035] The term "loudness" is used herein to refer to level (eg, average level) or power (eg, average power).

[0036] Thus, slow smoother 2 is configured to determine the average loudness of the input audio signal, where the average is over a time period longer (e.g., much longer) than the DRC application time (e.g., a typical DRC application time) implemented by subsystems 6 and 7. In some embodiments (e.g., the embodiment described below with reference to FIG. 2A ), the loudness is determined from metadata provided with (e.g., included in) the input audio signal. The output of slow smoother 2 is used by gain adjustment subsystem 4 to limit the gain output by DRC subsystem 6 (i.e., the gain indicated by the time-varying gain g(t) output from DRC subsystem 6). Subsystem 4 is coupled and configured to operate in response to a target (target level or power) identified in FIG. 2 as “target.” When the output of slow smoother 2 (average of input audio signal level or power) approaches (or meets or exceeds) the target, subsystem 4 asserts control data (identified in FIG. 2 as a "gain adjust" value) to subsystem 6 to cause the gain output by subsystem 6 to approach a gain of 1 (i.e., equal to a gain of 1, or closer to a gain of 1 than if subsystem 4 were disabled or omitted). In some implementations, the difference between the current output of slow smoother 2 and the target determines the current output of each of the gain adjust values.

[0037] Thus, the DRC system of FIG. 2 "gets out of the way" (i.e., applies a reduced DRC, e.g., no DRC) if the output of slow smoother 2 approaches (or meets or exceeds) the target, and otherwise applies full DRC (i.e., the DRC that would be applied if subsystems 2 and 4 were omitted or disabled).

[0038] The target (in response to which subsystem 4 operates) may be the knee point of the DRC applied by subsystem 6 (e.g., when the output of slow smoother 2 is at least substantially equal to the knee point, subsystem 6 outputs a gain of 1). The knee point may be the input signal level above which the DRC gain curve specifies a gain value less than 1; thus, when the output of slow smoother 2 is at least substantially equal to such a knee point, it is reasonable to assume that the input audio is already compressed. In some other embodiments, the target is a value equal to (or substantially equal to) the maximum playback level of the playback system or device that reproduces the output audio. At such a target, it is also reasonable to assume that the input audio is already compressed.

[0039] One traditional approach to DRC (e.g., the DRC performed by the system of FIG. 1) is to apply some kind of slow-moving automatic gain control (AGC) before the DRC, match the average level of the AGC-leveled audio content to a desired target level (before applying the DRC), and implement the DRC gain curve to have a guard band around this target level (so that for each segment of AGC-leveled audio whose average level falls within the guard band, the DRC is not applied; e.g., the DRC applies a gain of 1). However, when traditional DRC (with or without a guard band) is performed on AGC-leveled audio, undesirable DRC artifacts (e.g., pumping and / or breathing) can occur. Also, traditional guard bands may be suboptimally positioned to reduce pumping or other artifacts. Additionally, playback systems (implementing traditional DRC) often do not know the average mastering level of the original content (before the application of AGC), and therefore the system's behavior cannot reasonably assume that reduced DRC (or no DRC) should be applied to any segment of AGC-leveled audio. Depending on the characteristics of the original (pre-AGC) content, it may not be desirable to run DRC on AGC-leveled audio, even if there is a guard band around the AGC target level. Also, unknown digital volume control is often applied to the input audio, and traditional AGC levelers will "counter" such volume control, resulting in an unattractive user experience during playback.

[0040] In contrast to traditional approaches, the system of FIG. 2 performs DRC in a manner controlled (in part) by the average level or power of the input audio (i.e., the output of slow smoother 2 of FIG. 2), where the average is determined over an interval longer (e.g., much longer) than each attack time and / or each release time of the application (at non-unity gain) of DRC by subsystems 6 and 7 (or longer (e.g., much longer) than the duration of each application, including attack and release, by subsystems 6 and 7 of DRC at non-unity gain). When AGC is performed on the input audio signal of FIG. 2, the signal level (or power) whose average is determined (by slow smoother 2) (by subsystem 4 of FIG. 2) so that the resulting average level or power controls the DRC must be the level or power of the original signal before the application of AGC.

[0041] To understand the benefits of typical operation of the system of FIG. 2 (and other embodiments of the first class), consider the case where the input audio signal is a well-mastered signal that, on average, is precisely at the correct level, and thus the input audio signal (the signal undergoing DRC) has already passed through various compression processes. This is likely to be the case for music, since engineers know that mastered music is likely to be played in environments with a lot of background noise (such as in a car). The best that a DRC embodiment of the present invention can do in this case is to leave the input signal completely untouched. Even well-compressed tracks contain soft (quiet) bits. For example, a drum hit is typically followed by a reverberant tail that encompasses nearly the entire input level and should gradually fade away. Typically, any traditional DRC method results in boosting the quiet bits, since that is essentially what dynamic range compression means. However, the resulting output audio does not sound as good as doing nothing. Running traditional DRC on a well-mastered track that is already compressed would deviate from the artist's intent.

[0042] In contrast, the DRC control applied by the system of FIG. 2 (i.e., the control performed by elements 2 and 4 of FIG. 2) is determined by a slowly smoothed (averaged) version of the input signal level (or power), so that if the averaged input level or power (as indicated by the output of slow smoother 2) is close to (or matches or exceeds) the target (as would be the case for well-mastered, sufficiently loud input audio), the DRC is effectively disabled or reduced.

[0043] In practice, when audio is played back through a laptop or cell phone (or other playback system or device with limited power processing capabilities), the user typically struggles to obtain a playback level sufficient for comfortable listening. To achieve reasonable loudness playback using a playback device, the target digital average playback level typically must be very high. The exact value of the target average level is highly dependent on the particular device, but it is typically high (i.e., substantially equal to or greater than the typical value of the target provided to subsystem 4 in the embodiment of FIG. 2). For music (or other audio) to achieve such a high input average level, it must already have undergone significant compression. Thus, it turns out that the best thing for a DRC system to do is to leave it alone. The operation of a typical implementation of the system of FIG. 2 disables or reduces the application of DRC to such music (or other audio).

[0044] Input audio (e.g., a music track) with an average input level (determined by the output of a typical implementation of slow smoother 2 of FIG. 2) significantly lower than the target (provided to subsystem 4 of the embodiment of FIG. 2) will have peaks that are much louder than the average level and quieter bits that are lower than the average level. In this case, it should be assumed that the lower digital level is intended to be audible and that further dynamic range compression is required. The operation of a typical implementation of the system of FIG. 2 does not disable or reduce the application of DRC to such input audio.

[0045] It should be understood that the slow-moving average of the level or power of the audio signal (as indicated by the output of slow smoother 2 in FIG. 2) already imposes a constraint on the dynamic range of the signal. If the slow-moving average is close to the DRC knee point, or close to the maximum signal level, the DRC curve (implemented by DRC subsystem 6 in FIG. 2) no longer needs to be applied (or even applied at all). ("Maximum signal level" here refers to the maximum playback level of the intended playback system or device.)

[0046] Referring now to Figure 2A, another embodiment in the first class of embodiments above will be described. The system of Figure 2A (a dynamic range compressor or DRC) differs from the system of Figure 2 (also a dynamic range compressor or DRC) in that the DRC of Figure 2A includes an average loudness determination subsystem 8 rather than a slow smoother 2. The other elements of Figure 2A (elements 1, 4, 6, and 7) are identical to the corresponding (same-numbered) elements of Figure 2. Subsystem 6 of Figure 2A may be implemented in any of the ways that subsystem 6 of Figure 2 may be implemented (e.g., including elements 3 and 5 of Figure 1).

[0047] A DRC according to some embodiments of the present invention (e.g., the exemplary embodiment of FIG. 2A ) is implemented in a digital signal processor (e.g., programmed in software) configured to apply loudness-based processing to input audio data representing an audio program (e.g., implementing a “Dolby Volume” playback volume control or another type of playback loudness control or loudness leveling). Such loudness-based processing may use metadata (provided with and corresponding to the audio data) that indicates the loudness of the audio content and / or the loudness processing state of the audio content (e.g., what type of loudness processing has been performed on the audio content).

[0048] In the system of FIG. 2A , average loudness determination subsystem 8 is configured to determine (and provide to gain adjustment subsystem 4) data indicative of the average loudness (e.g., average level) of the input audio data (input audio signal), where the average is over a longer (e.g., much longer) time ("averaging" time) than each attack and / or release time of the dynamic range compression applied by subsystems 6 and 7 (or than the duration of each application, including attack and release, of DRC by subsystems 6 and 7 at a non-unity gain). In some implementations, the input audio data is included in a Dolby Digital bitstream (which also includes metadata corresponding to the audio data). In some implementations, subsystem 8 may be configured to parse the metadata to identify metadata indicative of the loudness of each segment (e.g., frame) of a sequence of segments of the input audio data and, if necessary, determine from the metadata a sequence of such loudness averages (where each average in the sequence is over a sufficiently long averaging time). In some implementations, subsystem 8 may implement models of middle ear and cochlear behavior, as well as a psychoacoustic model of loudness.

[0049] The output of subsystem 8 is used (in subsystem 4 of FIG. 2A) in the same way that the output of slow smoother 2 is used (in subsystem 4) in the embodiment of FIG.

[0050] 2A system (or subsystem 6 of another embodiment of the present invention) implements a DRC curve that maps input loudness to output loudness (thus subsystem 6 outputs loudness difference values ​​rather than gain values). In such implementations, subsystem 7 maps each loudness difference value to a corresponding gain change to be applied to the input audio signal.

[0051] For many applications, it may be desirable to employ a multi-band implementation of the system (dynamic range compressor) of FIG. 2 or FIG. 2A (or a multi-band implementation of the system of FIG. 3 or FIG. 4, described below).

[0052] In a multiband implementation of the system of FIG. 2 (or FIG. 2A), the input audio is divided into multiple frequency bands (e.g., by a filter bank). For each frequency band, an average loudness (e.g., average loudness, or average level, or average power, as indicated by metadata or averaged metadata) is determined (by subsystem 2 of FIG. 2 or subsystem 8 of FIG. 2A). For each band, the average loudness is measured over a longer (e.g., much longer) time period (the "averaging" time) than the attack and / or release times of the dynamic range compression applied by subsystems 6 and 7 (or the duration of each application, including attack and release, of DRC by subsystems 6 and 7 at non-unity gains). Subsystem 6 implements a set of DRC gain curves (one DRC gain curve for each frequency band). For each frequency band, a sequence of gain values ​​(output from subsystem 6 for that band) is applied (by subsystem 7) to the corresponding band of the input audio, thereby generating each frequency band of the output audio. The frequency bands of the output audio may be combined to generate the output audio signal. In other words, a DRC gain value (for each band) determined by the DRC gain curve is applied to each band of the input audio to generate "dynamic range compressed" audio in each band, and the "dynamic range compressed" audio (for each of the individual bands) can then be combined to form the output audio signal. In one embodiment, a DRC gain may thus be determined for at least one frequency band of a plurality of frequency bands, and the DRC gain may be applied to the frequency band.

[0053] In some implementations, subsystem 4 uses the average loudness values ​​determined for each of the individual frequency bands to determine the “gain adjustment” values ​​(shown in FIGS. 2 and 2A) used to control the application of DRC (in such implementations, each of the gain adjustment values ​​relates to an individual frequency band of the input audio). In some other implementations, a single (wideband) average loudness value is determined (by subsystem 2 of FIG. 2 or subsystem 8 of FIG. 2A), and this single average loudness value (which typically varies over time because it relates to a sequence of different segments of the input audio) is used by subsystem 4 to determine the “gain adjustment” values ​​(shown in FIGS. 2 and 2A) used to control the application of DRC. In the latter implementations, each of the gain adjustment values ​​(which relate to the wideband input audio, rather than to the individual frequency bands of the input audio) is applied to all of the DRC gain curves (and each DRC gain curve relates to a different one of the frequency bands). In some of the latter implementations, the difference between the wideband average loudness value and the target is used to generate the gain adjustment value used to control the application of the per-band DRC gain curve.

[0054] In a multi-band implementation, the DRC gain determination (applied by subsystem 7) typically involves smoothing the gains for each of the bands (e.g., the gains determined by the DRC gain curves for each of the bands) across the bands to improve timbre. In a multi-band implementation, different bands may have different DRC knee points, and thus the target (typically a selected broadband target) does not necessarily coincide with a specific knee point for each of the bands.

[0055] The inventors contemplate that various known methods that may reduce pumping and breathing artifacts in DRCs (e.g., some methods of the type implemented in "Dolby Volume" loudness level levelers) may be implemented in combination with some embodiments of the DRCs of the present invention. Examples of such methods include: Auditory scene analysis, where gain changes are applied with greater intensity in response to auditory scene changes; Hierarchical Constraints: The gains in individual frequency bands are constrained by the channel gains, which are in turn constrained by the sum level. For example, a DRC system implemented according to FIG. 2A (or FIG. 2) above may also implement audio scene analysis to further control the execution of the DRC.

[0056] A second class of embodiments of the present invention is directed to reducing pumping artifacts during DRC for input audio having regular transient components (e.g., a sequence of identical or similar transient components). Exemplary embodiments in the second class control (e.g., include a subsystem configured to control) a release time constant (of the release of each application of dynamic range compression). This includes implementing a first release time constant (referred to as a relatively slow release time constant) when a segment of the input audio signal contains regular transient components (including by applying a smoothed dynamic range compression gain to the segment of the input audio signal) and implementing a relatively fast release time constant (i.e., a release time constant faster than the first release time constant) when a different segment of the input audio signal does not contain regular transient components (including by applying an unsmoothed dynamic range compression gain to the different segment of the input audio signal). When the relatively slow release time constant is implemented, pumping artifacts are reduced or prevented from occurring.

[0057] In audio (especially music), there are often regular transients that cause a repeated attack and release of a typical DRC (dynamic range compressor). This can result in a well-known, unpleasant artifact (produced by dynamic range compression) known as pumping. Certain aspects of the present invention aim to solve this problem and provide technical advantages by modifying the release behavior of a dynamic range compressor.

[0058] A second class of exemplary embodiments will now be described with reference to Figure 3. The dynamic range compressor of Figure 3 differs from the conventional dynamic range compressor of Figure 1 in that the compressor of Figure 3 includes a smoother 11 and a gain adjustment subsystem 13 combined as shown (but not in the compressor of Figure 1). The other elements of the DRC of Figure 3 (elements 1, 3, 5, and 7) are identical to the corresponding (identically numbered) elements of the DRC of Figure 1. The DRC gain output from the DRC gain curve subsystem 5 is identified in Figure 3 as gain "gDRC."

[0059] Elements 3, 5, 11, and 13 of Figure 3 comprise a DRC gain determination subsystem (implemented in accordance with an embodiment of the present invention) that replaces conventional DRC gain determination subsystem 6 of Figure 1. In some variations of the implementation shown in Figure 3, elements 3 and 5 of Figure 3 are replaced by one of the above-described alternative implementations of DRC gain determination subsystem 6 of Figure 1.

[0060] 3 embodiment, DRC gain smoother 11 is provided to smooth the DRC gain (gDRC) output from DRC gain curve subsystem 5, thereby generating a smoothed DRC gain (identified as "gDRCsmoothed" in FIG. 3). The smoothed DRC gain and gain gDRC (output from subsystem 5) are provided to subsystem 13. In some implementations, gain adjustment subsystem 13 is configured to output the smaller of each gain value gDRC and the corresponding smoothed gain gDRCsmoothed (as the current gain value of g(t) applied by subsystem 7), such that the output of subsystem 13 in response to each gain value gDRC is: min(gDRC,gDRCsmoothed) This becomes:

[0061] When a segment of input audio has regular transients (e.g., a sequence of identical or similar transients, such as a sequence of drum hits), smoother 11 catches up with subsystem 5. This means that subsystem 13 reaches a state where it outputs (i.e., provides to subsystem 7) the current "gDRCsmoothed" value (output from smoother 11) rather than the corresponding gain value "gDRC" output from subsystem 5. During such operation, subsystem 7's application of the gDRCsmoothed value (rather than the corresponding value gDRC) effectively delays the release of the system's DRC application, thereby reducing (or preventing) pumping artifacts. In typical operation (in response to a segment of input audio with regular transients), subsystem 13 initially outputs the current gain value gDRC, which causes the system of FIG. 3 to operate in a state that provides some fast-release behavior. This continues until a state is reached where subsystem 13 outputs the current "gDRCsmoothed" value (rather than the corresponding gDRC value), at which point the system of Figure 3 implements a slower release (i.e., implements a relatively slow release time constant), thereby reducing or preventing the occurrence of pumping artifacts. In typical operation, (in response to a segment of input audio that does not have regular transients) subsystem 13 outputs the current gain value gDRC (thus implementing a relatively fast release time constant) rather than the current "gDRCsmoothed" value.

[0062] We have found it useful to implement subsystem 13 to operate in response to a user-specified parameter p. Selecting different values ​​for parameter p (sometimes referred to as the "pumping parameter") allows the user to trade off pumping artifacts and loudness. In such an implementation, subsystem 13 outputs a final gain g (i.e., one value of the time-varying gain g(t)) in response to each gain value gDRC and the corresponding smoothed gain gDRCsmoothed. The final gain g is the value: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with a user-selectable value ranging from 0 to 1.

[0063] Thus, if a user selects p to be equal to (or approximately equal to) 1, the average loudness of the output audio may be increased (compared to the average output audio loudness when p=0), but undesirable pumping artifacts may occur. If a user selects p to be equal to (or approximately equal to) 0, the average loudness of the output audio may be decreased (compared to the average output audio loudness when p=1), but the occurrence of pumping artifacts may be reduced or prevented.

[0064] In a preferred embodiment, the system of Figure 3 is implemented as a multi-band compressor. In such an implementation, the value of gDRCsmooothed, and typically also the value of gDRC, is determined on a band-by-band basis, allowing for different choices of the pumping parameter p for different frequency bands. It may be useful to allow lower frequency bands to have larger values ​​of p.

[0065] According to some implementations of the present invention, a DRC system belongs to both the first class of embodiments and the second class of embodiments. For example, a system may implement artifact reduction aspects of the second class of embodiments (e.g., its DRC gain determination subsystem may include elements 11 and 13 of the implementation of FIG. 3) and DRC reduction aspects of the first class of embodiments (e.g., it may include elements identical to or corresponding to slow smoother 2 and subsystem 4 of FIG. 2, or elements corresponding to subsystems 8 and 4 of FIG. 2A).

[0066] A third class of embodiments of the present invention is directed to reducing breathing artifacts during DRC for decaying input audio. Exemplary embodiments of the third class control (e.g., include a subsystem configured to control) a release time constant (of the release of each application of dynamic range compression) in response to a loudness gradient of the input audio signal. This control typically implements a faster release time constant in response to increased steepness of the loudness gradient (to reduce or prevent the occurrence of breathing artifacts) and a slower release time constant in response to decreased steepness of the loudness gradient (to reduce or prevent the occurrence of pumping artifacts).

[0067] A third class of exemplary embodiments will now be described with reference to Figure 4. The dynamic range compressor of Figure 4 differs from the conventional dynamic range compressor of Figure 3 in that the compressor of Figure 4 includes a coupled loudness gradient estimation subsystem 15 as shown. The other elements of the DRC of Figure 4 (elements 1, 3, 5, 7, 11, and 13) are identical to the corresponding (same-numbered) elements of the DRC of Figure 3.

[0068] Elements 3, 5, 11, 13, and 15 of Figure 4 comprise a DRC gain determination subsystem (implemented in accordance with an embodiment of the present invention) that can replace conventional DRC gain determination subsystem 6 of Figure 1. In a variation of the implementation shown in Figure 4, elements 3 and 5 of Figure 4 are replaced by one of the above-described alternative implementations of DRC gain determination subsystem 6 of Figure 1.

[0069] Breathing artifacts are well-known artifacts that can result from dynamic range compression and are particularly annoying when the input audio becomes quieter (attenuates) and the DRC (dynamic range compressor) applies increasing gain to it (for example, during the release period of the dynamic range compression application). Depending on the relative time constants of the decaying input audio and the compressor release, breathing artifacts can increase the loudness of the output audio when the listener (or audio content creator) is expecting the audio to be getting quieter.

[0070] The average loudness (level or power) of an input audio signal typically varies over time and has a gradient (sometimes referred to herein as a loudness gradient). The gradient is the rate of change of the averaged level or power of the input audio signal over time. In this context, the time period over which the average loudness is determined need not be longer (or much longer) than the DRC application time described above. According to one aspect of the embodiment of FIG. 4 , subsystem 15 is provided to generate a loudness gradient estimate (e.g., a time-smoothed estimate) of the average loudness (mean level or power) of the input audio signal. In the implementation of FIG. 4 , subsystem 15 is configured to generate this loudness gradient estimate based on the estimated level or power (of the input audio signal) determined by level estimator 1. Alternatively, the loudness gradient estimate is generated in another manner (e.g., based on loudness metadata corresponding to the input audio).

[0071] The estimated level or power (of the input audio signal) determined by subsystem 1 typically varies over time, and subsystem 15 may be configured to determine, for each time period, a time-smoothed estimate of the loudness gradient (from the corresponding sequence of estimated levels or powers output from subsystem 1). In response to the loudness gradient estimate, subsystem 15 generates a control signal (identified as “control” in FIG. 4 ) and provides the control signal to smoother 11. In response to increasing steepness of the loudness gradient (i.e., increasing values ​​of positive loudness gradient or increasing values ​​of negative loudness gradient (or less negative values)), the control signal generated by subsystem 15 varies the time constant of the smoothing performed by smoother 11, allowing for a faster release time constant (of the release of each application of dynamic range compression by the system of FIG. 4 ). In other words, in response to increasing steepness of the loudness gradient, the control signal generated by subsystem 15 varies the time constant of the smoothing performed by smoother 11, such that the smoothed gain value (gDRCsmoothed) output from smoother 11 causes subsystem 13 to output a gain value that effectively allows a faster release of the dynamic range compression application. The faster release time constant resulting from increasing steepness of the loudness gradient typically reduces (or prevents) breathing artifacts.

[0072] In response to the decreasing steepness of the loudness gradient, the control signal generated by subsystem 15 varies the time constant of the smoothing performed by smoother 11, allowing the release time constant (of each application of dynamic range compression by the system of FIG. 4) to be slower. As discussed above with reference to FIG. 3, such a slower release time constant can reduce (or prevent) pumping artifacts and can also reduce (or prevent or make less noticeable) breathing artifacts.

[0073] In a preferred embodiment, the time constant used (by smoother 11) to calculate the value gDRCsmoothed (in response to the value gDRC) is scaled by the loudness gradient (determined by subsystem 15 from level estimates generated on the full wideband input audio) to be in the range of about 2 seconds to about 6 seconds.

[0074] In a variation of the system of Figure 4 (which is an alternative embodiment of the present invention), elements 3 and 5 of Figure 4 are replaced by an implementation of DRC gain determination subsystem 6 (e.g., any of the implementations of subsystem 6 of Figure 2), and the control signal generated by loudness gradient estimation subsystem 15 is used to directly control (e.g., increase) the release time of such subsystem 6 rather than to control smoother 11. In such an embodiment, elements 11 and 13 are optionally omitted.

[0075] According to some embodiments of the present invention, a DRC system belongs to both the first class of embodiments and the third class of embodiments. For example, a system may implement both the artifact reduction aspects of the third class of embodiments (e.g., its DRC gain determination subsystem may include elements 11, 13, and 15 of the implementation of FIG. 4) and the DRC reduction aspects of the first class of embodiments (e.g., may include elements identical to or corresponding to slow smoother 2 and subsystem 4 of FIG. 2, or elements corresponding to subsystems 8 and 4 of FIG. 2A).

[0076] An example embodiment (EE) of the present invention includes: EE1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, comprising: (a) determining an average loudness of the input audio signal, said average over a time period longer than a DRC application time of a DRC, said DRC application time being an attack time or a release time of an instance of application of the DRC, or a duration of an instance of application of the DRC; (b) when the average loudness of the input audio signal approaches, meets, or exceeds a target, applying a reduced DRC to the input audio signal to generate the output audio signal, and otherwise applying a full DRC to the input audio signal to generate the output audio signal. method. EE2. The method of EE1, wherein the target is an audio signal level at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device that reproduces the output audio signal. EE3. The method of EE1 or EE2, wherein the input audio signal has a plurality of frequency bands, and step (b) includes determining a DRC gain for each of the frequency bands and applying the DRC gain to each of the frequency bands. EE4. The method of EE3, wherein step (a) includes determining a wideband average loudness of the input audio signal, and step (b) includes applying a reduced DRC to each frequency band when the wideband average loudness approaches, matches, or exceeds a target. EE5. The method of EE3, wherein step (a) includes determining an average loudness for each of the frequency bands, and step (b) includes applying the reduced DRC to each of the frequency bands whose average loudness approaches, meets, or exceeds the target. EE6. The method of EE3, wherein determining a DRC gain includes smoothing gains for individual frequency bands of the individual frequency bands across individual frequency bands of the frequency bands to improve timbre. EE7. The method of EE1, EE2, EE3, EE4, EE5, or EE6, wherein step (b) comprises: Determine the dynamic DRC gain gDRC; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; applying the dynamic gain g to an input audio signal; method. EE8. The method of EE7, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with a value in the range of 0 to 1. EE9. The method of EE1, EE2, EE3, EE4, EE5, EE6, EE7, or EE8, wherein the input audio signal has a loudness gradient, the method further comprising: controlling release time constants for application of the reduced DRC and the full DRC in response to a loudness gradient of the input audio signal; method. EE10. The method of EE9, wherein the release time constant is controlled to be faster in response to increased steepness of the loudness gradient and to be slower in response to decreased steepness of the loudness gradient. EE11. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, the method comprising: determining a level estimate of the input audio signal; determining a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; applying the dynamic gain g to an input audio signal to thereby generate the output audio signal. method. EE12. The method of claim EE11, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with a value in the range of 0 to 1. EE13. The input audio signal has a loudness gradient, and the method further comprises: controlling a release time constant for application of the DRC to the input audio signal in response to a loudness gradient of the input audio signal. The method according to EE11 or EE12. EE14. The method of EE13, wherein the release time constant is controlled to be faster in response to increased steepness of the loudness gradient and to be slower in response to decreased steepness of the loudness gradient. EE15. The method of EE13, wherein controlling the release time constant includes controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmooothed. EE16. A method according to any one of claims 1 to 5, wherein the input audio signal has a plurality of frequency bands, and wherein the dynamic gain g comprises individual band gains for each one of the frequency bands, and wherein applying the dynamic gain g comprises: applying individual band gains to individual bands of the frequency bands of the input audio signal; method. EE17. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a level estimation subsystem coupled and configured to determine a level estimate of the input audio signal; a DRC gain curve subsystem coupled and configured to determine a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate; a gain determination subsystem coupled and configured to smooth the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed, and determine a dynamic gain g, including by determining a minimum value of each pair of corresponding values ​​of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; a gain application subsystem coupled and configured to apply the dynamic gain g to an input audio signal to generate the output audio signal; the gain determination subsystem is configured to determine the dynamic gain g such that when applying the dynamic gain g to a segment of the input audio signal that includes regular transient components, the system implements a first release time constant, and when applying the dynamic gain g to a different segment of the input audio signal that does not include regular transient components, the system implements a release time constant that is faster than the first release time constant. system. EE18. The system of EE17, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with selectable values ​​in the range 0 to 1 for the system. EE19. The input audio signal has a loudness gradient, and the gain determination subsystem is configured to control a release time constant for application of a DRC to the input audio signal in response to the loudness gradient of the input audio signal. EE17 or EE18. EE20. The system of EE19, wherein the gain determination subsystem is configured to provide a faster release time constant in response to increased steepness of the loudness gradient and a slower release time constant in response to decreased steepness of the loudness gradient. EE21. The method of EE19, wherein the gain determination subsystem is configured to control a release time constant, including by controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmooothed. EE22. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a loudness determination subsystem coupled and configured to determine an average loudness of the input audio signal, the average over a time period longer than a DRC application time of a DRC, the DRC application time being an attack time or a release time of an instance of application of the DRC, or a duration of an instance of application of the DRC; a gain determination and application subsystem coupled and configured to apply a reduced DRC to the input audio signal when the average loudness of the input audio signal approaches, meets, or exceeds a target, thereby generating the output audio signal, and otherwise apply a full DRC to the input audio signal to generate the output audio signal. system. EE23. The system of EE22, wherein the target is an audio signal level at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device that reproduces the output audio signal. EE24. The system of EE22 or EE23, wherein the input audio signal has a plurality of frequency bands, and wherein the gain determination and application subsystem is configured to determine a DRC gain for each one of the frequency bands and apply the DRC gain to the each one of the frequency bands. EE25. The gain determination and application subsystem: Determine the dynamic DRC gain gDRC; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; configured to apply the dynamic gain g to an input audio signal. EE22, EE23 or EE24. EE26. The system of claim EE25, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is the pumping parameter with a value ranging from 0 to 1 for the system. EE27. The system of EE22, EE23, EE24, EE25, or EE26, wherein the input audio signal has a loudness gradient, and the gain determination and application subsystem: configured to control release time constants for application of the reduced DRC and the full DRC in response to a loudness gradient of the input audio signal. system. EE28. The system of EE22, EE23, EE24, EE25, EE26, or EE27, wherein the gain determination and application subsystem is configured to make the release time constant faster in response to increased steepness of the loudness gradient and slower in response to decreased steepness of the loudness gradient. Various modifications to the implementations described in this disclosure will be readily apparent to those skilled in the art. The general principles defined herein may be applied to other implementations without departing from the spirit or scope of the disclosure. Accordingly, the claims are not intended to be limited to the specific implementations described and illustrated herein, but are to be accorded the widest scope consistent with this disclosure.

[0077] The methods and systems described in this disclosure may be implemented as software, firmware, and / or hardware. For example, certain components (e.g., elements 1, 2, 4, 6, and 7 of FIG. 2, or elements 1, 4, 6, 7, and 8 of FIG. 2A, or elements 1, 3, 5, 7, 11, and 13 of FIG. 3, or elements 1, 3, 5, 7, 11, 13, and 15 of FIG. 4) may be implemented as software running on a digital signal processor (e.g., having an input coupled to receive an input audio signal) or a microprocessor. Some components may be implemented as hardware and / or as application-specific integrated circuits. Signals encountered in the methods and systems described above may be stored in media such as random access memory or optical storage media. They may be transmitted over a network, such as an airwave network, a satellite network, a wireless network, or a wired network such as the Internet. Typical devices that utilize the methods and systems described in this disclosure are portable electronic devices or other consumer equipment used to store and process (e.g., implement playback or render) audio signals (e.g., output audio signals generated according to any embodiment of the systems or methods of the present invention).

[0078] In this application, a description in the form "A and / or B" means "A" or "B" or "A and B."

[0079] Several aspects will be described. [Aspect 1] 1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, comprising: (a) determining an average loudness of the input audio signal, loudness representing a level or power of the input audio signal, the average being over a time longer than a DRC application time of the DRC, the DRC application time being the attack time or release time of an instance of application of the DRC, or the duration of an instance of application of the DRC including the attack and release, that applies a non-unity gain after an attack and before a release, or that a DRC gain curve determines a non-unity gain; (b) applying a reduced DRC to the input audio signal to generate the output audio signal when the average loudness of the input audio signal approaches, meets, or exceeds a target, and otherwise applying a full DRC to the input audio signal to generate the output audio signal, wherein the target is an audio signal level at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device that reproduces the output audio signal. method. [Aspect 2] 2. The method of claim 1, wherein the input audio signal has multiple frequency bands. Aspect 3 3. The method of claim 2, wherein step (b) comprises determining a DRC gain for at least one frequency band of the plurality of frequency bands and applying the DRC gain to the frequency band. Aspect 4 3. The method of claim 2, wherein step (b) comprises determining DRC gains for the respective ones of the frequency bands and applying those DRC gains to the respective ones of the frequency bands. Aspect 5 5. The method of claim 4, wherein determining the DRC gains includes smoothing gains for individual frequency bands of the frequency bands across individual frequency bands of the frequency bands to improve timbre. Aspect 6 3. The method of claim 2, wherein step (a) comprises determining a wideband average loudness of the input audio signal, and step (b) comprises applying the reduced DRC to each frequency band when the wideband average loudness approaches, matches, or exceeds a target. Aspect 7 3. The method of embodiment 2, wherein step (a) comprises determining an average loudness for each of the frequency bands, and wherein step (b) comprises applying the reduced DRC to each frequency band whose average loudness approaches, meets, or exceeds the target. Aspect 8 Step (b) is: Determine the dynamic DRC gain gDRC; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; applying the dynamic gain g to the input audio signal. 8. The method of any one of embodiments 1 to 7. Aspect 9 9. The method of claim 8, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with a value in the range of 0 to 1. Aspect 10 10. The method of embodiment 9, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable. Aspect 11 The input audio signal has a loudness gradient, and the method further comprises: controlling release time constants for application of the reduced DRC and full DRC in response to the loudness gradient of the input audio signal. 11. The method of any one of embodiments 1 to 10. Aspect 12 12. The method of claim 11, wherein the release time constant is controlled to be faster in response to increased steepness of the loudness gradient and controlled to be slower in response to decreased steepness of the loudness gradient. Aspect 13 1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, the method comprising: determining a level estimate of the input audio signal; determining a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; applying the dynamic gain g to the input audio signal to thereby generate the output audio signal. method. Aspect 14 14. The method of claim 13, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with a value in the range of 0 to 1. Aspect 15 15. The method of embodiment 14, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable. Aspect 16 The input audio signal has a loudness gradient, and the method further comprises: controlling a release time constant for application of a DRC to the input audio signal in response to the loudness gradient of the input audio signal. 16. The method of any one of embodiments 13 to 15. Aspect 17 17. The method of embodiment 16, wherein the release time constant is controlled to be faster in response to increased steepness of the loudness gradient and controlled to be slower in response to decreased steepness of the loudness gradient. Aspect 18 Aspect 16. The method of aspect 16, wherein controlling the release time constant includes controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmooothed. Aspect 19 19. The method of any one of aspects 13 to 18, wherein the input audio signal has a plurality of frequency bands, and the dynamic gain g comprises individual band gains for respective ones of the frequency bands, and applying the dynamic gain g comprises: applying the individual band gains to individual bands of the frequency bands of the input audio signal. method. Aspect 20 1. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a level estimation subsystem coupled and configured to determine a level estimate of the input audio signal; a DRC gain curve subsystem coupled and configured to determine a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate; a gain determination subsystem coupled and configured to smooth the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed, and determine a dynamic gain g, including by determining a minimum value of each pair of corresponding values ​​of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; a gain application subsystem coupled and configured to apply the dynamic gain g to the input audio signal, thereby generating the output audio signal; the gain determination subsystem is configured to determine the dynamic gain g such that when applying the dynamic gain g to a segment of the input audio signal that includes regular transient components, the system implements a first release time constant, and when applying the dynamic gain g to a different segment of the input audio signal that does not include regular transient components, the system implements a release time constant that is faster than the first release time constant. system. Aspect 21 21. The system of claim 20, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter with selectable values ​​in the range 0 to 1 for the system. Aspect 22 22. The system of embodiment 21, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable. Aspect 23 the input audio signal has a loudness gradient, and the gain determination subsystem is configured to control a release time constant for application of a DRC to the input audio signal in response to the loudness gradient of the input audio signal. The system according to embodiment 20 or embodiment 22. Aspect 24 24. The system of aspect 23, wherein the gain determination subsystem is configured to make the release time constant faster in response to increased steepness of the loudness gradient and to make the release time constant slower in response to decreased steepness of the loudness gradient. Aspect 25 24. The system of claim 23, wherein the gain determination subsystem is configured to control the release time constant, including by controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmooothed. Aspect 26 1. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a loudness determination subsystem coupled and configured to determine an average loudness of the input audio signal, wherein loudness represents a level or power of the input audio signal, the average being over a time period longer than a DRC application time of the DRC, the DRC application time being an attack time or a release time of an instance of application of the DRC, or a duration of an instance of application of the DRC including an attack and a release, that applies a non-unity gain after an attack and before a release, or where a DRC gain curve determines a non-unity gain; a gain determination and application subsystem coupled and configured to apply a reduced DRC to the input audio signal to generate the output audio signal when the average loudness of the input audio signal approaches, meets, or exceeds a target, and otherwise apply a full DRC to the input audio signal to generate the output audio signal, wherein the target is an audio signal level at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device that reproduces the output audio signal; system. Aspect 27 27. The system of embodiment 26, wherein the input audio signal has multiple frequency bands. Aspect 28 28. The system of aspect 27, wherein the gain determination and application subsystem is configured to determine a DRC gain for at least one frequency band among the plurality of frequency bands and apply the DRC gain to the frequency band. Aspect 29 28. The system of aspect 27, wherein the gain determination and application subsystem is configured to determine DRC gains for individual ones of the frequency bands and apply those DRC gains to the individual ones of the frequency bands. Aspect 30 30. The system of claim 29, wherein determining the DRC gain includes smoothing gains for individual frequency bands of the frequency band across individual frequency bands of the frequency band to improve timbre. Aspect 31 28. The system of aspect 27, wherein the gain determination and application subsystem is configured to determine a wideband average loudness of the input audio signal and apply the reduced DRC to each frequency band when the wideband average loudness approaches, matches, or exceeds the target. Aspect 32 28. The system of claim 27, wherein the gain determination and application subsystem determines an average loudness for each of the frequency bands and applies the reduced DRC to each frequency band whose average loudness approaches, matches, or exceeds the target. Aspect 33 The gain determination and application subsystem: Determine the dynamic DRC gain gDRC; smoothing the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on a minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; configured to apply the dynamic gain g to the input audio signal. 33. The system of any one of aspects 26 to 32. Aspect 34 34. The system of claim 33, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is the pumping parameter with a value ranging from 0 to 1 for the system. Aspect 35 35. The system of embodiment 34, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable. Aspect 36 The input audio signal has a loudness gradient, and the gain determination and application subsystem: configured to control release time constants for application of the reduced DRC and full DRC in response to the loudness gradient of the input audio signal. A system described in any one of aspects 26 to 35. Aspect 37 37. The system of embodiment 36, wherein the gain determination and application subsystem is configured to make the release time constant faster in response to increased steepness of the loudness gradient and slower in response to decreased steepness of the loudness gradient. Aspect 38 37. The system of embodiment 36, wherein controlling the release time constant includes controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmoothed.

Claims

1. 1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, comprising: (a) determining an average loudness of the input audio signal, where loudness represents a level or power of the input audio signal, and the average loudness is determined over a time period that is longer than a DRC application time of the DRC, the DRC application time being the attack time or release time of an instance of application of the DRC that applies a non-unity gain or where a DRC gain curve determines a non-unity gain after an attack time and before a release time, or the duration of the instance of application of the DRC including the attack time and the release time; (b) applying a reduced DRC to the input audio signal to generate the output audio signal by controlling a DRC gain to approach a gain of 1 when the average loudness of the input audio signal approaches, matches, or exceeds a target, and otherwise applying the DRC gain to the input audio signal to generate the output audio signal, wherein the target is a knee point for the DRC. method.

2. The method of claim 1 , wherein the input audio signal comprises multiple frequency bands.

3. The method of claim 2 , wherein step (b) includes determining a DRC gain for at least one frequency band of the plurality of frequency bands and applying the DRC gain to the frequency band.

4. 3. The method of claim 2, wherein step (b) comprises determining DRC gains for the respective ones of the frequency bands and applying those DRC gains to the respective ones of the frequency bands.

5. 5. The method of claim 4, wherein determining the DRC gains includes smoothing gains for individual frequency bands of the frequency bands across individual frequency bands of the frequency bands to improve timbre.

6. 3. The method of claim 2, wherein step (a) comprises determining a wideband average loudness of the input audio signal, and step (b) comprises applying the reduced DRC to each frequency band when the wideband average loudness approaches, meets, or exceeds a target.

7. 3. The method of claim 2, wherein step (a) comprises determining a respective average loudness for each of the frequency bands, and step (b) comprises applying the reduced DRC to each frequency band whose average loudness approaches, meets, or exceeds the target.

8. Step (b) is: determining a DRC gain gDRC; smoothing the DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; generating a dynamic gain g, where if the DRC gain gDRC is less than the smoothed dynamic gain gDRCsmoothed, the dynamic gain g is the DRC gain gDRC, otherwise the dynamic gain g is a weighted average of gDRC and gDRCsmoothed; applying the dynamic gain g to the input audio signal. The method of claim 1.

9. 9. The method of claim 8, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter having a value in the range of 0 to 1.

10. 10. The method of claim 9, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable.

11. The input audio signal has a loudness gradient, and the method further comprises: controlling release time constants for application of the reduced DRC and full DRC in response to the loudness gradient of the input audio signal. The method of claim 1.

12. 12. The method of claim 11, wherein the release time constant is controlled to be faster in response to an increased steepness of the loudness gradient and is controlled to be slower in response to a decreased steepness of the loudness gradient.

13. 1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, the method comprising: determining a level estimate of the input audio signal; determining a DRC gain gDRC by applying a DRC gain curve to the level estimate; smoothing the DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; generating a dynamic gain g, where if the DRC gain gDRC is less than the smoothed dynamic gain gDRCsmoothed, the dynamic gain g is the DRC gain gDRC, otherwise the dynamic gain g is a weighted average of gDRC and gDRCsmoothed; applying the dynamic gain g to the input audio signal to thereby generate the output audio signal. method.

14. 14. The method of claim 13, wherein the dynamic gain g is: g=p*gDRC+(1-p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter having a value in the range of 0 to 1.

15. 15. The method of claim 14, wherein the value of the pumping parameter in the range of 0 to 1 is user selectable.

16. The input audio signal has a loudness gradient, and the method further comprises: controlling a release time constant for application of a DRC to the input audio signal in response to the loudness gradient of the input audio signal. The method of claim 13.

17. 17. The method of claim 16, wherein the release time constant is controlled to be faster in response to an increased steepness of the loudness gradient and to be slower in response to a decreased steepness of the loudness gradient.

18. Claim 16. The method of claim 16, wherein controlling the release time constant comprises controlling a time constant for performing smoothing to generate the smoothed dynamic gain gDRCsmooothed.

19. 14. The method of claim 13, wherein the input audio signal has a plurality of frequency bands, and the dynamic gain g comprises individual band gains for each one of the frequency bands, and applying the dynamic gain g comprises: applying the individual band gains to individual bands of the frequency bands of the input audio signal. method.

20. 1. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a level estimation subsystem coupled and configured to determine a level estimate of the input audio signal; a DRC gain curve subsystem coupled and configured to determine a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate; a gain determination subsystem coupled and configured to smooth the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed, and determine a dynamic gain g, including by determining the minimum of each pair of corresponding values ​​of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; a gain application subsystem coupled and configured to apply the dynamic gain g to the input audio signal, thereby generating the output audio signal; the gain determination subsystem is configured to determine the dynamic gain g such that when applying the dynamic gain g to a segment of the input audio signal that includes regular transient components, the system implements a first release time constant, and when applying the dynamic gain g to a different segment of the input audio signal that does not include regular transient components, the system implements a release time constant that is faster than the first release time constant. system.

21. 1. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a loudness determination subsystem coupled and configured to determine an average loudness of the input audio signal, wherein loudness represents a level or power of the input audio signal, and wherein the average loudness is determined over a time period that is longer than a DRC application time of the DRC, the DRC application time being an attack time or a release time of an instance of application of the DRC that applies a non-unity gain or a DRC gain curve determines a non-unity gain after an attack time and before a release time, or a duration of an instance of application of the DRC that includes the attack time and the release time; a gain determination and application subsystem coupled and configured to apply a reduced DRC to the input audio signal to generate the output audio signal by controlling a DRC gain to approach a gain of 1 when the average loudness of the input audio signal approaches, meets, or exceeds a target, and to otherwise apply the DRC gain to the input audio signal to generate the output audio signal, wherein the target is a knee point for the DRC. system.

22. 22. The system of claim 21, wherein the input audio signal comprises multiple frequency bands.

23. 23. The system of claim 22, wherein the gain determination and application subsystem is configured to determine a DRC gain for at least one frequency band of the plurality of frequency bands and apply the DRC gain to the frequency band.

24. 23. The system of claim 22, wherein the gain determination and application subsystem is configured to determine DRC gains for the respective ones of the frequency bands and apply those DRC gains to the respective ones of the frequency bands.

25. 25. The system of claim 24, wherein determining the DRC gains includes smoothing gains for individual frequency bands of the frequency bands across individual frequency bands of the frequency bands to improve timbre.

26. 23. The system of claim 22, wherein the gain determination and application subsystem is configured to determine a wideband average loudness of the input audio signal and apply the reduced DRC to each frequency band when the wideband average loudness approaches, meets, or exceeds the target.

27. 23. The system of claim 22, wherein the gain determination and application subsystem includes determining a respective average loudness for each of the frequency bands and applying the reduced DRC to each frequency band whose average loudness approaches, meets, or exceeds the target.

Citation Information

Patent Citations

  • Calculation and adjustment of perceived loudness and / or perceived spectral balance of audio signals

    JP2008518565A

  • Audio compressor

    JP2012065068A

  • Method for compressing the dynamics in an audio signal

    US20160381468A1

  • Dynamic range control apparatus

    WO2013038451A1