Dynamic Range Compression with Reduced Artifacts

The dynamic range compression system addresses the issue of unwanted artifacts in audio signals by applying reduced DRC based on average loudness and controlling release time constants, resulting in improved listenability and reduced artifacts in systems with limited power processing.

JP7682860B2Active Publication Date: 2025-05-26DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022516103
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-09-10
Publication Date
2025-05-26
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

Existing dynamic range compression (DRC) systems often introduce unwanted artifacts such as 'pumping' and 'breathing' when applied to audio signals, especially in systems with limited power processing capabilities.

Method used

The proposed method involves a dynamic range compression system that applies reduced or no DRC when the average loudness of the input audio signal approaches or exceeds a target level, thereby minimizing artifacts. This system also controls the release time constant based on the presence of regular transients and the loudness gradient to further reduce artifacts.

Benefits of technology

The method effectively reduces or prevents the occurrence of pumping and breathing artifacts, while maximizing average loudness and maintaining the listenability of audio content, even in systems with limited power processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682860000001
    Figure 0007682860000001
  • Figure 0007682860000002
    Figure 0007682860000002
  • Figure 0007682860000003
    Figure 0007682860000003
Patent Text Reader

Abstract

A method for generating output audio for playback by a system or device with limited power processing capabilities, preferably performing dynamic range compression (DRC) on the audio in a manner intended to reduce or prevent undesirable artifacts (e.g., pumping and / or breathing) in the output audio. Some embodiments perform DRC to maximize average loudness during playback (while preventing the loss of quieter elements) and reduce or prevent distortion. Another aspect is a system or device configured to perform embodiments of the method. In some embodiments, if the average loudness of the input audio approaches (or meets or exceeds) a target (e.g., a knee point for DRC, or a signal level near the maximum playback level of the intended playback system), reduced DRC is applied because such input audio is assumed to be already compressed. Otherwise, full DRC is applied to the input audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to European Patent Application No. 19197154.8, filed on September 13, 2019, and U.S. Provisional Patent Application No. 62 / 899,769, filed on September 13, 2019, which are hereby incorporated by reference in their entirety.

[0002] Technical Field The present disclosure relates to dynamic range compression of audio signals.

Background Art

[0003] Here, dynamic range compression may be referred to as "DRC", and a dynamic range compressor may also be referred to as "DRC".

[0004] As shown in FIG. 1, a traditional dynamic range compressor ("DRC") includes a level estimator 1, a dynamic range compression (DRC) gain determination subsystem 6, and a gain application subsystem 7 coupled as shown. In some embodiments, the gain determination subsystem 6 includes a smoother 3 and a dynamic range compression (DRC) gain curve subsystem 5 coupled as shown. The subsystem 7 is coupled and configured to apply a time - varying gain g(t) (e.g., a series of gain values) output from the subsystem 6 (e.g., from the subsystem 5 of the subsystem 6 as shown in FIG. 1) to an input audio signal to generate an audio output signal.

[0005] The dynamic range compression applied to the input audio signal reduces the power of segments of the input audio signal where the smoothed level (output from smoother 3) exceeds a threshold (the first knee point), i.e., subsystem 7 applies a segment gain (s) determined by subsystem 6 that is less than gain 1, and may increase the power of segments of the input audio signal where the smoothed level (output from smoother 3) is below a second threshold (the "lower" knee point), i.e., subsystem 7 may apply a gain greater than 1 determined by subsystem 6 to such segments.

[0006] Dynamic range compression (DRC) has an attack (starting at or near each point in time when the input audio level (e.g., indicated by the output of level estimator 1) rises to the knee point, having a duration known as the attack time, an attack interval) and a release (starting at or near each point in time when the input audio level drops to the knee point, having a duration known as the release time, a release interval). Full DRC is applied during the interval between after the attack and before the release. The amount of DRC applied can increase from zero (when subsystem 6 outputs gain 1) to the full amount during the attack and then return to zero during the release.

[0007] In DRC, one important possibility is that there is a large gain for low-level input audio below the knee point, and then the gain monotonically decreases to gain 1 as the input signal level approaches its maximum level. Thus, DRC does not actually reduce the high levels of the input audio, but rather only increases the lower levels of the input audio. Such a case occurs when DRC is used in an attempt to obtain maximum loudness from a low-power system.

[0008] When the input audio signal jumps to its maximum or minimum level, this change is indicated by the output of the level estimator 1 in FIG. 1. However, the DRC gain (output from subsystem 6) typically does not immediately change to its full value (i.e., the value it would have if the input audio signal level were fixed at the maximum or minimum level and had not jumped to that level). Instead, the DRC gain changes smoothly to its full value (e.g., due to the presence of the smoother 3 in FIG. 1). The time required for the DRC gain to reach its full value (after the input signal level has jumped to the maximum level) is the attack time, and the time required for the DRC gain to reach its full value (after the input signal level has jumped to the minimum level) is the release time. Since the DRC gain typically approaches its full value exponentially, when attack and release times are cited (e.g., in system specifications), the cited attack (or release) time is often the time required to reach the midpoint as the DRC gain approaches its full value.

[0009] In the implementation shown in FIG. 1, the smoother 3 (i.e., the smoothing time constant used by the smoother 3) determines the attack time and the release time (the attack time may be equal to, different from, or one or both may be zero).

[0010] In other implementations of subsystem 6 (not specifically shown in FIG. 1), the smoother 3 is replaced by another element (e.g., a smoother that acts on the gain value output from the DRC gain curve subsystem 5) or a subsystem that determines the attack and release times for each section of DRC application by subsystem 6. Generally, subsystem 6 is configured to determine the attack time and the release time for each section (i.e., each section where subsystem 6 outputs a gain other than 1) where subsystem 6 applies DRC (e.g., in response to user selection).

[0011] In some implementations of dynamic range compression, the performance of compression is split, with the gain application subsystem (e.g., subsystem 7) implemented in the decoder or playback system or device, other elements of the compression (e.g., subsystems 1 and 6) implemented in the encoder, and the gain g(t) being sent as metadata (to the decoder or playback system or device) along with the input audio in the encoded bitstream. Some embodiments of the present invention (described below) contemplate being implemented in this way.

[0012] The level estimator 1 is coupled and configured to determine a level estimate value and provide it to subsystem 6 (e.g., to the smoother 3 in the implementation shown in FIG. 1). The level estimate value is an estimate of the loudness of the input audio signal (typically changing over time) (e.g., the level estimate value represents a sequence of average levels or average power values, and the averaging time for each is long enough for the stability of the dynamic range compression applied by the system of FIG. 1). One typical level estimate value is the average power. Another example of a level estimate value is the loudness defined by the ITU-R BS.1770 loudness standard. The smoother 3 is coupled and configured to apply smoothing to the level estimate value output from the estimator 1 and generate a smoothed estimate value of the average level or power of the input audio signal (smoothed level estimate value) (and assert it to subsystem 5).

[0013] In the implementation shown in FIG. 1, in response to the smoothed level estimate value (e.g., a smoothed sequence of average level or power values) determined by the smoother 3, the subsystem 5 determines a sequence of values of the gain g(t). The subsystem 5 implements a function (typically called a "DRC gain curve") that maps each value output from the smoother 3 (of the average level or power) to a value of the gain g(t) (gain value). The gain element 7 applies the gain g(t) to the input audio signal to generate an output audio signal (which is a dynamically range compressed version of the input audio signal). This is, for example, by applying each value of the sequence of values of g (the gain) to the corresponding value of the input audio signal (e.g., the sequence of its values).

[0014] In some other implementations, the subsystem 6 applies a DRC gain curve that maps each value output from the level estimator 1 (not from the smoother 3 as shown in FIG. 1) (of the average level or power) to a value of the gain g(t) (gain value) (including by implementing DRC attack and release (with determined attack and release times for each section of the DRC application by the subsystem 6)), and the gain value g(t) (optionally modified to implement an attack or release) is provided to the gain element 7. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0015] Some embodiments of the present invention are directed to a method of performing dynamic range compression (DRC) on an audio signal in a manner that generates output audio for playback by a system or device having limited power processing capabilities (e.g., a notebook, laptop, tablet, soundbar, mobile phone, or other device for use with or for use with a small speaker), preferably reducing or preventing the occurrence of unwanted artifacts (e.g., those known as "pumping" and "breathing") in the output audio. In some embodiments, the DRC maximizes (or provides a sufficiently large) average loudness during playback (while preventing loss of quieter elements of the audio) and also reduces or prevents distortion (e.g., to reduce or prevent the occurrence of pumping and / or breathing artifacts and / or to reduce or prevent tonal variations due to frequency components generated by the non-linear gain application of the DRC). Some embodiments perform DRC on an audio signal in a manner intended to optimize the general listenability of content for radio broadcasts or audio content or components within an audio stream.

[0016] Here, the expression "DRC application time" may be used to indicate the attack time (or release time) of an instance of DRC application (e.g., an instance of application that applies a gain other than 1, or an instance of application where the DRC gain curve determines a gain other than 1 after attack and before release), or the duration of such an instance of DRC application (including attack and release).

[0017] In a first class of embodiments of the dynamic range compression (DRC) method of the present invention, when the average loudness [size] (e.g., average level or power) of the input audio signal approaches (or matches or exceeds) the target, a reduced DRC gain (e.g., no DRC) is applied to the input audio signal. This is because such an input audio signal is assumed to already be compressed (e.g., maximizing loudness while preventing loss of quieter elements of the audio during playback). Otherwise, full DRC is applied to the input audio signal. The application of reduced DRC (e.g., no DRC) when full DRC is not necessary reduces or prevents the occurrence of pumping and / or breathing artifacts that would result from full DRC. The average loudness is determined over a time (the "averaging" time) that is longer (e.g., much longer) than the DRC application time of the DRC. The target may be a target signal level or target signal power. In a typical embodiment of the first class, the target is the knee point for the DRC, or an audio signal level that is close to (e.g., equal to or substantially equal to) the maximum playback level of the playback system or device that plays the output audio.

[0018] A second class of embodiments of the dynamic range compression (DRC) method of the present invention is directed to reducing pumping artifacts during the execution of DRC on input audio having regular transients (e.g., a sequence of identical or similar transients). A typical embodiment of the second class controls the release time constant (of the release of each application of dynamic range compression). This includes implementing a first release time constant (referred to as a relatively slow release time constant) when a segment of the input audio signal contains regular transients (including by applying a smoothed dynamic range compression gain to that segment of the input audio signal), and implementing a relatively fast release time constant (i.e., a release time constant faster than the first release time constant) when a different segment of the input audio signal does not contain regular transients (including by applying an unsmoothed dynamic range compression gain to that different segment of the input audio signal). When a relatively slow release time constant is implemented, pumping artifacts are reduced or their occurrence is prevented.

[0019] A third class of embodiments of the dynamic range compression (DRC) method of the present invention is directed to reducing breathing artifacts during the execution of DRC on decaying input audio. A typical embodiment of the third class controls the release time constant (of the release of each application of dynamic range compression) in response to the loudness gradient of the input audio signal. This control typically implements a faster release time constant (to reduce or prevent the occurrence of breathing artifacts) in response to an increased steepness of the loudness gradient, and implements a slower release time constant (to reduce or prevent the occurrence of pumping artifacts) in response to a decreased steepness of the loudness gradient.

[0020] Another aspect of the present invention is a system (e.g., a dynamic range compressor) or apparatus configured to perform any embodiment of the method of the present invention on an input audio signal. In certain classes of embodiments, the present invention performs dynamic range compression (in accordance with any embodiment of the method of the present invention) to produce dynamically compressed audio and is configured to perform playback of the dynamically compressed audio (e.g., a notebook, laptop, tablet, sound bar, cellular phone, or other apparatus having (or for use with) small speakers, or a playback system having a limited (e.g., physically limited) power processing capacity).

[0021] In some embodiments, the system of the present invention is a general-purpose or special-purpose processor programmed (or otherwise configured) in software (or firmware) to perform an embodiment of the method of the present invention and / or includes the same. In some embodiments, the system of the present invention is a general-purpose processor coupled to receive input audio data and programmed (with appropriate software) to produce output audio data by performing an embodiment of the method of the present invention. In some embodiments, the system of the present invention is a digital signal processor coupled to receive input audio data and configured (e.g., programmed) to produce output audio data in response to the input audio data by performing an embodiment of the method of the present invention.

[0022] Aspects of the present invention include a system (e.g., programmed) configured to perform any embodiment of the method of the present invention and a computer-readable medium (e.g., a disk) storing code for performing any embodiment of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023]

Figure 1

[0024]

Figure 2

[0025]

Figure 2A

[0026]

Figure 3

[0027]

Figure 4

DETAILED DESCRIPTION OF THE INVENTION

[0028] Some embodiments of the present invention provide improvements and technical solutions for reducing or preventing the occurrence of undesirable (e.g., annoying) artifacts known as "pumping" and "breathing" as a result of dynamic range compression. Different classes of embodiments implement different approaches (described herein) for preventing or reducing such pumping and breathing artifacts.

[0029] A first class of embodiments of the dynamic range compressor and dynamic range compression (DRC) method of the present invention will be described with reference to FIG. 2. In some embodiments of this class, DRC applies (or is configured to apply) reduced DRC (or no DRC) to the input audio signal when the average level (or power) of the input audio signal approaches (or equals or exceeds) a target (e.g., a high average level). This is because such an input audio signal is assumed to already be compressed. The average level (or power) of the input audio signal is determined over a time (“averaging time”) that is longer (e.g., much longer) than each attack time and / or each release time of the dynamic range compression, and the reduced DRC is applied “when” the signal has the average level (or power). This is in the sense that it is applied to each segment (having a duration greater than the averaging time) of the input audio signal having the average level (or power). Referring to the first class of embodiments, the expression “target” is used broadly to indicate a target signal level (or signal power), the value of which is such that it can be reasonably assumed that an input audio signal whose average level (or power) approaches (or equals or exceeds) the target is already compressed. In typical embodiments of the first class, the target level is an audio signal level close to the knee point for DRC or the maximum playback level of a playback system or device that plays the output audio.

[0030] The dynamic range compressor of FIG. 2 is an exemplary embodiment of the first class of embodiments. Unlike the conventional dynamic range compressor of FIG. 1, the compressor of FIG. 2 includes a slow smoother 2 coupled as shown to a gain adjustment subsystem 4 (the compressor of FIG. 1 does not). The other elements (elements 1, 6, and 7) of the DRC of FIG. 2 are the same as the corresponding (same numbered) elements of the DRC of FIG. 1. Subsystem 6 of FIG. 2 can be implemented in any of the ways in which subsystem 6 of FIG. 1 can be implemented (e.g., including elements 3 and 5 of FIG. 1).

[0031] A typical DRC gain curve (e.g., implemented by subsystem 5 of FIG. 1 or FIG. 2) determines gain values (values of gain g(t) of FIG. 1 or FIG. 2) as follows: In response to the audio signal having a high average level or power (e.g., the smoothed level output from smoother 3) above a threshold (knee point), less than 1 In response to the audio signal having an average level or power (e.g., the smoothed level output from smoother 3) below the knee point but above a second threshold (second knee point), 1, and In response to the audio signal having a low average level or power below the second threshold, greater than 1, Thereby, the DRC (which generates an output audio signal in response to an input audio signal) maintains a monotonic increase in the average output audio signal level (or power) as the average input signal level (or power) increases. The gain values determined by the DRC gain curve (and the output from subsystem 6 of FIG. 1 or FIG. 2) vary dynamically in response to changes in the average level (or power) of the input audio signal. However, this can cause undesirable artifacts (e.g., pumping artifacts) in the output audio signal, particularly when the average level (or power) of the input signal is near the knee point of the DRC gain curve.

[0032] The inventors have come to recognize that, if the input audio signal is already sufficiently compressed (and can be reproduced at sufficient loudness by the intended playback device), an ideal DRC system would apply less DRC (e.g., no DRC) to such an input audio signal than to other input audio signals. In other words, an ideal DRC system would be "non-intrusive" (and thus not introduce significant pumping artifacts) for a sufficiently compressed (and sufficiently loud) input audio signal. The inventors have further recognized that, if the input audio signal has an average level or power that approaches (or matches or exceeds) an appropriate target value (e.g., a target value that is sufficiently high such that it can be reasonably assumed that the input audio signal is already compressed and can be reproduced at sufficient loudness by the intended playback device), and the average is determined over a sufficiently long time interval (i.e., a period that is longer (e.g., much longer) than each attack time and / or each release time of the dynamic range compression), the DRC system should apply less DRC to such an input audio signal than to other input audio signals (i.e., the DRC system should be "non-intrusive" and thus not introduce significant pumping artifacts into the input audio signal).

[0033] Referring again to FIG. 2, the slow smoother 2 generates (and provides to the gain adjustment subsystem 4) a slowly smoothed version of the level (or power) estimate output from the level estimator 1. The output of the slow smoother 2 is an estimate of the average level or power of the input audio signal, where the average is over a time that is longer (e.g., much longer) than the attack time (and / or release time) employed by the subsystem 6 (or the duration of each application of DRC with a gain other than 1 including attack and release by subsystems 6 and 7).

[0034] Here, the expression "DRC application time" is used to indicate the attack time (or the release time) of an instance of DRC application (e.g., an instance of application that applies a non - unity gain, or an instance of application where the DRC gain curve determines a non - unity gain after attack and before release), or the duration of such an instance of DRC application (including attack and release).

[0035] Here, the term "loudness" is used to denote level (e.g., average level) or power (e.g., average power).

[0036] Thus, the slow smoother 2 is configured to determine the average loudness of the input audio signal, where the average is over a time that is longer (e.g., much longer) than the DRC application time (e.g., a typical DRC application time) implemented by subsystems 6 and 7. In some embodiments (e.g., the embodiments described below with reference to FIG. 2A), the loudness is determined from metadata provided with (e.g., included in) the input audio signal. The output of the slow smoother 2 is used by the gain adjustment subsystem 4 to limit the gain output by the DRC subsystem 6 (i.e., the gain indicated by the time - varying gain g(t) output from the DRC subsystem 6). The subsystem 4 is coupled and configured to operate in response to a target (target level or power) identified as "target" in FIG. 2. As the output of the slow smoother 2 (the average of the input audio signal level or power) approaches (or equals or exceeds) the target, the subsystem 4 asserts control data (identified as "gain adjustment" value in FIG. 2) to the subsystem 6 and causes the gain output by the subsystem 6 to approach a gain of 1 (i.e., to equal the gain 1, or to be closer to the gain 1 than if the subsystem 4 were disabled or omitted). In some implementations, the difference between the current output of the slow smoother 2 and the target determines each current output of the gain adjustment value.

[0037] Thus, the DRC system of FIG. 2 "makes itself unobtrusive" (i.e., applies reduced DRC, e.g., no DRC) if the output of the slow smoother 2 approaches (or matches, or exceeds) the target, and applies full DRC (i.e., the DRC that would be applied if subsystems 2 and 4 were omitted or disabled) otherwise.

[0038] The target (in response to which subsystem 4 operates) may be the knee point of the DRC applied by subsystem 6 (e.g., if the output of the slow smoother 2 is at least substantially equal to the knee point, subsystem 6 outputs a gain of 1). The knee point may be an input signal level above which the DRC gain curve specifies a gain value less than 1, and thus it is reasonable to assume that the input audio is already compressed if the output of the slow smoother 2 is at least substantially equal to such a knee point. In some other embodiments, the target is a value equal to (or substantially equal to) the maximum playback level of a playback system or device that plays the output audio. With such a target, it is also reasonable to assume that the input audio is already compressed.

[0039] One traditional approach to DRC (e.g., DRC executed by the system of FIG. 1) is to apply some slow-acting AGC (automatic gain control) before the DRC, match the average level of the AGC-leveled audio content to a desired target level (before applying the DRC), and implement the DRC gain curve to have a guard band around this target level (thereby, for each segment of the AGC-leveled audio having an average level within the guard band, the DRC is not applied, e.g., the DRC applies a gain of 1). However, when traditional DRC (with or without a guard band) is executed on AGC-leveled audio, undesirable DRC artifacts (e.g., pumping and / or breathing) may occur. Also, conventional guard bands may be placed in a suboptimal way to reduce pumping or other artifacts. Further, in a playback system (implementing traditional DRC), the average mastering level of the original content (before AGC application) is often unknown, and thus, it cannot be reasonably assumed that reduced DRC (or no DRC) should be applied to any segment of the AGC-leveled audio. Depending on the characteristics of the original (pre-AGC) content, it may not be desirable to execute DRC on the AGC-leveled audio even if there is a guard band around the AGC target level. Also, unknown digital volume controls are often applied to the input audio, and traditional AGC leveling devices "counteract" such volume controls, resulting in an unappealing user experience during playback.

[0040] In contrast to traditional approaches, the system of FIG. 2 performs DRC in a manner controlled by the average level or power of the input audio (i.e., the output of the slow smoother 2 of FIG. 2), where the average is longer (e.g., much longer) than each attack time and / or each release time of the application of DRC (at a non - unity gain) by subsystems 6 and 7 (or longer (e.g., much longer) than the duration of each application, including attack and release, of DRC by subsystems 6 and 7 at a non - unity gain) and is determined over an interval. When AGC is performed on the input audio signal of FIG. 2, the resulting average value level or power for controlling DRC (by subsystem 4 of FIG. 2) and for which the average is determined (by the slow smoother 2) must be the level or power of the original signal before the application of AGC.

[0041] To understand the advantages of the typical operation of the system of FIG. 2 (and other embodiments of the first class), consider the case where the input audio signal is a well - mastered signal that is, on average, exactly at the correct level, and thus the input audio signal (the signal that undergoes DRC) has already passed through various compressions. This is likely to be the case for music. This is because engineers know that mastered music is likely to be played in an environment with a lot of background noise (such as in a car). The best thing that the DRC embodiments of the present invention can do in this case is to leave the input signal completely as it is. Even highly compressed tracks contain soft (quiet) bits. For example, after a drum hit, there generally follows a reverberation tail, which should contain almost all input levels and fade away gradually. Typically, any traditional DRC method will result in boosting the quiet bits. This is essentially what dynamic range compression means. However, the resulting output audio will not sound as good as doing nothing. Performing traditional DRC on a well - mastered track that is already compressed will deviate from the artist's intent.

[0042] In contrast, the DRC control applied by the system of FIG. 2 (i.e., the control executed by elements 2 and 4 of FIG. 2) is determined by a slowly smoothed (averaged) version of the input signal level (or power), such that if the averaged input level or power (indicated by the output of the slow smoother 2) is close to (or matches, or exceeds) the target (would be the case for sufficiently mastered, sufficiently loud input audio), the DRC is effectively disabled or reduced.

[0043] In practice, when audio is played back by a laptop or mobile phone (or other playback system or device with limited power processing capabilities), typically the user struggles to obtain a playback level sufficient to listen comfortably. To achieve proper loudness playback using the playback device, typically the target digital averaged playback level has to be very high. The exact value of the target averaged level depends very much on the particular device, but is typically high (i.e., substantially equal to or greater than the typical value of the target provided to subsystem 4 of the embodiment of FIG. 2). For music (or other audio) to achieve such a high input average level, it already has to have undergone a fair amount of compression. Thus, it can be seen that the best thing for the DRC system to do is to leave it alone. The operation of a typical implementation of the system of FIG. 2 disables or reduces the application of DRC to such music (or other audio).

[0044] Input audio (e.g., a music track) having an average input level significantly lower (determined by the output of a typical implementation of the slower smoother 2 in the embodiment of FIG. 2) than the target (provided to subsystem 4 of the embodiment of FIG. 2) has peaks that are much louder than the average level and quiet bits that are lower than the average level. In this case, it is intended that lower digital levels be audible, and it should be assumed that further dynamic range compression is needed. The operation of a typical implementation of the system of FIG. 2 does not disable or reduce the application of DRC to such input audio.

[0045] It should be understood that the slow-moving average of the level or power of the audio signal (indicated by the output of the slower smoother 2 in FIG. 2) already places a constraint on the dynamic range of the signal. When the slow-moving average is close to the DRC knee point or close to the maximum signal level, the DRC curve (implemented by the DRC subsystem 6 in FIG. 2) need no longer be applied (or need not be applied fully) (the "maximum signal level" here indicates the maximum playback level of the intended playback system or device).

[0046] Next, referring to FIG. 2A, another embodiment in the first class of the above-described embodiments will be described. The system of FIG. 2A (dynamic range compressor or DRC) differs from the system of FIG. 2 (also a dynamic range compressor or DRC) in that the DRC of FIG. 2A includes an average loudness determination subsystem 8 instead of the slower smoother 2. The other elements of FIG. 2A (elements 1, 4, 6, and 7) are the same as the corresponding (same-numbered) elements of FIG. 2. The subsystem 6 of FIG. 2A can be implemented in any of the ways in which the subsystem 6 of FIG. 2 can be implemented (e.g., including elements 3 and 5 of FIG. 1).

[0047] The DRC according to some embodiments of the present invention (e.g., the exemplary embodiment of FIG. 2A) is implemented in a digital signal processor (e.g., one implementing "Dolby Volume" playback volume control, or another type of playback loudness control or loudness normalization) configured (e.g., programmed in software) to apply loudness-based processing to input audio data representing an audio program. Such loudness-based processing can use metadata (provided with the audio data and corresponding to the audio data) indicating the loudness of the audio content and / or the loudness processing state of the audio content (e.g., what type of loudness processing has been performed on the audio content).

[0048] In the system of FIG. 2A, the average loudness determination subsystem 8 is configured to determine (and provide to the gain adjustment subsystem 4) data indicating the average loudness (e.g., average level) of the input audio data (input audio signal), where the average is over a time (the "averaging" time) longer than (or, in the case of DRC applied by subsystems 6 and 7 with a gain other than 1, longer than the duration of each application, including attack and release) each attack time and / or each release time of the dynamic range compression applied by subsystems 6 and 7. In some implementations, the input audio data is included in a "Dolby Digital" bitstream (including metadata corresponding to the audio data). In some embodiments, subsystem 8 may be configured to parse the metadata to identify metadata indicating the loudness of each segment (e.g., frame) of a sequence of segments of the input audio data, and, if necessary, determine from the metadata a sequence of such loudness averages (where each average in the sequence is over a sufficiently long averaging time). In some implementations, subsystem 8 may implement a model of the behavior of the middle ear and cochlea, as well as a psychoacoustic model of loudness.

[0049] The output of subsystem 8 is used (in subsystem 4 of FIG. 2A) in the same way that the output of the slow smoother 2 is used (in subsystem 4) in the embodiment of FIG. 2.

[0050] In some alternative implementations, subsystem 6 of the system of FIG. 2A (or subsystem 6 of another embodiment of the present invention) implements a DRC curve that maps the input loudness to the output loudness (thus, subsystem 6 outputs a loudness difference value rather than a gain value). In such an implementation, subsystem 7 maps each loudness difference value to the corresponding gain change to be applied to the input audio signal.

[0051] For many applications, it is considered desirable to employ a multi-band implementation of the system of FIG. 2 or FIG. 2A (dynamic range compressor) (or a multi-band implementation of the system of FIG. 3 or FIG. 4 described later).

[0052] In the multi-band implementation of the system of FIG. 2 (or FIG. 2A), the input audio is split into a plurality of frequency bands (e.g., by a filter bank). For each of the frequency bands, the average loudness (e.g., the average loudness indicated by metadata or averaged metadata, or average level, or average power) is determined (by subsystem 2 of FIG. 2 or subsystem 8 of FIG. 2A). For each band, the average loudness is over a time (the "averaging" time) that is longer (e.g., much longer) than each attack time and / or each release time of the dynamic range compression applied by subsystems 6 and 7 (or longer than the duration of each application of DRC by subsystems 6 and 7 with a gain other than 1, including attack and release). Subsystem 6 implements a set of DRC gain curves (one DRC gain curve for each frequency band). For each individual frequency band, a sequence of gain values (output from subsystem 6 for that band) is applied (by subsystem 7) to the corresponding band of the input audio, thereby generating each frequency band of the output audio. The frequency bands of the output audio may be combined to generate the output audio signal. In other words, the DRC gain values (for each band) determined by the DRC gain curves are applied to the individual bands of the input audio to generate "dynamically range compressed" audio in each band, and then the "dynamically range compressed" audio (in each of the individual bands) can be combined to form the output audio signal. In some embodiments, the DRC gain may be determined for at least one frequency band of the plurality of frequency bands and the DRC gain may be applied to the frequency band.

[0053] In some implementations, subsystem 4 uses the average loudness value determined for each of the individual frequency bands to determine the "gain adjustment" value (shown in FIGS. 2 and 2A) used to control the application of DRC (in such implementations, each of the gain adjustment values relates to an individual frequency band of the input audio). In some other implementations, a single (broadband) average loudness value is determined (by subsystem 2 in FIG. 2 or subsystem 8 in FIG. 2A), and this single average loudness value (which typically varies over time since it relates to a sequence of different segments of the input audio) is used by subsystem 4 to determine the "gain adjustment" value (shown in FIGS. 2 and 2A) used to control the application of DRC. In the latter implementations, each of the gain adjustment values (relating to broadband input audio rather than individual frequency bands of the input audio) is applied to all of the DRC gain curves (and each DRC gain curve relates to a different one of the frequency bands). In some of the latter implementations, the difference between the broadband average loudness value and the target is used to generate the gain adjustment value used to control the application of the per-band DRC gain curves.

[0054] In a multiband implementation, the determination of the DRC gain (applied by subsystem 7) typically includes smoothing the gains for the individual bands across the bands (e.g., the gains determined by the DRC gain curves for the individual bands across the bands) to improve sound quality. In a multiband implementation, different bands may have different DRC knee points, and thus the target (typically a selected broadband target) does not necessarily match the specific knee points for the individual bands.

[0055] The inventors consider that various known methods capable of reducing pumping and breathing artifacts in DRC (for example, some methods of the type implemented in a "Dolby Volume" loudness level normalizer) may be implemented in combination with some embodiments of the DRC of the present invention. Examples of such methods include the following: Auditory scene analysis. Gain changes are applied with greater intensity in response to auditory scene changes; Hierarchical constraints. The gain in an individual frequency band is constrained by the channel gain, and the channel gain is constrained by the total level. For example, a DRC system implemented according to FIG. 2A (or FIG. 2) described above can also implement audio scene analysis to further control the execution of DRC.

[0056] A second class of embodiments of the present invention is directed to reducing pumping artifacts during the execution of DRC for input audio having regular transient components (for example, a sequence of identical or similar transient components). A typical embodiment in the second class controls the release time constant (for example, includes a subsystem configured to control) of (the release of each application of dynamic range compression). This includes implementing a first release time constant (referred to as a relatively slow release time constant) when a segment of the input audio signal includes regular transient components (including by applying a smoothed dynamic range compression gain to that segment of the input audio signal), and implementing a relatively fast release time constant (i.e., a release time constant faster than the first release time constant) when different segments of the input audio signal do not include regular transient components (including by applying an unsmoothed dynamic range compression gain to the different segments of the input audio signal). When a relatively slow release time constant is implemented, pumping artifacts are reduced or their occurrence is prevented.

[0057] In audio (especially music), there are often regular transient components that cause the repeated attack and release of a normal DRC (Dynamic Range Compressor). This can result in a well-known unpleasant artifact, known as pumping, which is generated by dynamic range compression. One aspect of the present invention aims to solve this problem and provide a technical advantage by modifying the release behavior of a dynamic range compressor.

[0058] Referring to FIG. 3, an exemplary embodiment of a second class will be described. The compressor of FIG. 3 is different from the conventional dynamic range compressor of FIG. 1 in that a smoother 11 and a gain adjustment subsystem 13 are coupled as shown in the figure (but not so in the compressor of FIG. 1). Other elements (elements 1, 3, 5, and 7) of the DRC of FIG. 3 are the same as the corresponding (same-numbered) elements of the DRC of FIG. 1. The DRC gain output from the DRC gain curve subsystem 5 is identified as gain "gDRC" in FIG. 3.

[0059] The elements 3, 5, 11, and 13 of FIG. 3 comprise a DRC gain determination subsystem (implemented according to an embodiment of the present invention) that replaces the conventional DRC gain determination subsystem 6 of FIG. 1. In some variations of the implementation shown in FIG. 3, the elements 3 and 5 of FIG. 3 are replaced by one of the above-described alternative implementations of the DRC gain determination subsystem 6 of FIG. 1.

[0060] In the embodiment of FIG. 3, the DRC gain smoother 11 is provided to smooth the DRC gain (gDRC) output from the DRC gain curve subsystem 5, thereby generating a smoothed DRC gain (identified as "gDRCsmoothed" in FIG. 3). The smoothed DRC gain and the gain gDRC (output from subsystem 5) are provided to subsystem 13. In some implementations, the gain adjustment subsystem 13 is configured to output the smaller of each gain value gDRC and the corresponding smoothed gain gDRCsmoothed (as the current gain value of g(t) applied by subsystem 7), whereby in response to each gain value gDRC, the output of subsystem 13 is: min(gDRC,gDRCsmoothed) as follows.

[0061] If the segment of the input audio has regular transient components (e.g., a sequence of the same or similar transient components such as a sequence of drum hits), the smoother 11 catches up with the subsystem 5. This means that the subsystem 13 reaches a state where it outputs the current "gDRCsmoothed" value (i.e., provides it to the subsystem 7) instead of the corresponding gain value "gDRC" output from the subsystem 5. During such an operation, the application of the gDRCsmoothed value (instead of the corresponding value gDRC) by the subsystem 7 effectively delays the release of the DRC application by the system, thus reducing (or preventing the occurrence of) pumping artifacts. In typical operation (in response to a segment of input audio having regular transient components), the subsystem 13 initially outputs the current gain value gDRC, which operates the system in FIG. 3 in a state that provides some fast release behavior. This continues until the subsystem 13 reaches a state where it outputs the current "gDRCsmoothed" value (instead of the corresponding gDRC value), at which point the system in FIG. 3 implements a slower release (i.e., implements a relatively slow release time constant), thereby reducing or preventing the occurrence of pumping artifacts. In typical operation, (in response to a segment of input audio having no regular transient components) the subsystem 13 outputs the current gain value gDRC instead of the current "gDRCsmoothed" value (thus the system implements a relatively fast release time constant).

[0062] It has been found useful to implement the subsystem 13 to operate in response to a user-specified parameter p. By selecting different values of the parameter p (sometimes referred to as the "pumping parameter"), it enables the user to trade off pumping artifacts and loudness. In such an implementation, the subsystem 13 outputs a final gain g (i.e., one value of the time-varying gain g(t)) in response to each gain value gDRC and the corresponding smoothed gain gDRCsmoothed. The final gain g is the value: g = p * gDRC+(1 - p)*min(gDRC, gDRCsmoothed) where "p" is a pumping parameter having a value selectable by the user in the range from 0 to 1.

[0063] Thus, when the user selects p to be equal (or approximately equal) to 1, the average loudness of the output audio can be increased (compared to the average output audio loudness when p = 0), but undesirable pumping artifacts can occur. When the user selects p to be equal (or approximately equal) to 0, the average loudness of the output audio can be reduced (compared to the average output audio loudness when p = 1), but the occurrence of pumping artifacts can be reduced or prevented.

[0064] In a preferred embodiment, the system of FIG. 3 is implemented as a multiband compressor. In such an implementation, the value gDRCsmooothed, and typically also the value gDRC, are determined for each band, and different selections of the pumping parameter p can be made for different frequency bands. It may be useful to allow lower frequency bands to have larger p values of p.

[0065] According to some implementations of the present invention, the DRC system belongs to both a first class of embodiments and a second class of embodiments. For example, the system may implement the artifact reduction aspect of the second class of embodiments (e.g., its DRC gain determination subsystem may include elements 11 and 13 of the implementation of FIG. 3) and the DRC reduction aspect of the first class of embodiments (e.g., it may include elements identical or corresponding to the slow smoother 2 and subsystem 4 of FIG. 2, or elements corresponding to subsystem 8 and 4 of FIG. 2A).

[0066] A third class of embodiments of the present invention is directed to reducing breathing artifacts during the execution of DRC on attenuating input audio. A typical embodiment of the third class controls (e.g., includes a subsystem configured to control) the release time constant (of each application of dynamic range compression) in response to the loudness gradient of the input audio signal. This control typically implements a faster release time constant (to reduce or prevent the occurrence of breathing artifacts) in response to an increased steepness of the loudness gradient, and implements a slower release time constant (to reduce or prevent the occurrence of pumping artifacts) in response to a decreased steepness of the loudness gradient.

[0067] Referring to FIG. 4, an exemplary embodiment of the third class will be described. The dynamic range compressor of FIG. 4 differs from the conventional dynamic range compressor of FIG. 3 in that it includes a loudness gradient estimation subsystem 15 coupled as shown in the figure. The other elements (elements 1, 3, 5, 7, 11, and 13) of the DRC of FIG. 4 are the same as the corresponding (same numbered) elements of the DRC of FIG. 3.

[0068] Elements 3, 5, 11, 13, and 15 of FIG. 4 include a DRC gain determination subsystem (implemented in accordance with an embodiment of the present invention) that can replace the conventional DRC gain determination subsystem 6 of FIG. 1. In a variation of the implementation shown in FIG. 4, elements 3 and 5 of FIG. 4 are replaced by one of the above-described alternative implementations of the DRC gain determination subsystem 6 of FIG. 1.

[0069] Breathing artifact is a well-known artifact that may occur as a result of dynamic range compression, and is particularly troublesome when the input audio becomes quieter (attenuates) and the DRC (dynamic range compressor) applies an increasing gain to it (e.g., during the release interval of the application of dynamic range compression). Depending on the relative time constants of the attenuating input audio and the compressor release, the breathing artifact may increase the loudness of the output audio when the listener (or audio content creator) expects the audio to be getting quieter.

[0070] The averaged loudness (level or power) of the input audio signal typically changes over time and has a gradient (sometimes referred to in this paper as the loudness gradient). The gradient is the rate of change of the averaged level or power of the input audio signal over time. In this context, the time at which the averaged loudness is determined need not be longer (or much longer) than the above-described DRC application time. According to one aspect of the embodiment of FIG. 4, the subsystem 15 is provided to generate an estimate of the loudness gradient (e.g., a time-smoothed estimate) of the averaged loudness (average level or power) of the input audio signal. In the implementation of FIG. 4, the subsystem 15 is configured to generate this loudness gradient estimate based on the estimated level or power (of the input audio signal) determined by the level estimator 1. Alternatively, the loudness gradient estimate may be generated in another way (e.g., based on loudness metadata corresponding to the input audio).

[0071] The estimated level or power (of the input audio signal) determined by subsystem 1 typically changes over time, and subsystem 15 can be configured to determine, for each time, a time-smoothed estimate of the loudness gradient from the corresponding sequence of the estimated level or power output from subsystem 1. In response to the estimated value of the loudness gradient, subsystem 15 generates a control signal (identified as "Control" in FIG. 4) and provides the control signal to smoother 11. In response to the increasing steepness of the loudness gradient (i.e., an increasing value of the positive loudness gradient or an increasing value of the negative loudness gradient (or a value that is less negative)), the control signal generated by subsystem 15 changes the smoothing time constant executed by smoother 11 and allows for a faster release time constant for each application of dynamic range compression by the system of FIG. 4. In other words, in response to the increasing steepness of the loudness gradient, the control signal generated by subsystem 15 changes the smoothing time constant executed by smoother 11, whereby subsystem 13 outputs a gain value that effectively allows for a faster release of the dynamic range compression application by the gain value (gDRCsmoothed) output from smoother 11. The faster release time constant resulting from the increasing steepness of the loudness gradient typically reduces (or prevents the occurrence of) breathing artifacts.

[0072] In response to the decreasing steepness of the loudness gradient, the control signal generated by subsystem 15 changes the smoothing time constant executed by smoother 11 and allows for a slower release time constant for each application of dynamic range compression by the system of FIG. 4. As described above with reference to FIG. 3, such a slower release time constant can reduce (or prevent the occurrence of) pumping artifacts and can also reduce (or prevent the occurrence of or make less prominent) breathing artifacts.

[0073] In a preferred embodiment, the time constant used (by smoother 11) to calculate value gDRCsmoothed (in response to value gDRC) is scaled by a loudness gradient (determined by subsystem 15 from level estimates generated on a full broadband input audio) and is in the range of about 2 seconds to about 6 seconds.

[0074] In a variation of the system of FIG. 4 (which is an alternative embodiment of the present invention), elements 3 and 5 of FIG. 4 are replaced by an implementation of DRC gain determination subsystem 6 (e.g., any of the implementations of subsystem 6 of FIG. 2), and the control signal generated by loudness gradient estimation subsystem 15 is used not to control smoother 11, but to directly control (e.g., increase) the release time of such a subsystem 6. In such an embodiment, elements 11 and 13 are optionally omitted.

[0075] According to some embodiments of the present invention, the DRC system belongs to both the first class of embodiments and the third class of embodiments. For example, the system can implement both the artifact reduction aspect of the third class of embodiments (e.g., its DRC gain determination subsystem can include elements 11, 13, and 15 of the FIG. 4 implementation) and the DRC reduction aspect of the first class of embodiments (e.g., elements identical or corresponding to the slow smoother 2 and subsystem 4 of FIG. 2, or elements corresponding to subsystems 8 and 4 of FIG. 2A).

[0076] Exemplary embodiments (EE) of the present invention include: EE1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, comprising: (a) Determining an average loudness of an input audio signal, wherein the average is over a time longer than a DRC application time of DRC, and the DRC application time is an attack time or a release time of an instance of application of DRC, or a duration of an instance of application of DRC; (b) Applying a reduced DRC to the input audio signal when the average loudness of the input audio signal approaches, matches, or exceeds a target, thereby generating the output audio signal, or otherwise applying a full DRC to the input audio signal to generate the output audio signal; A method. EE2. The method according to EE1, wherein the target is an audio signal level that is at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device that plays the output audio signal. EE3. The method according to EE1 or EE2, wherein the input audio signal has a plurality of frequency bands, and step (b) includes determining a DRC gain for each of the individual frequency bands and applying the DRC gain to the individual frequency bands. EE4. The method according to EE3, wherein step (a) includes determining a broadband average loudness of the input audio signal, and step (b) includes applying a reduced DRC to each frequency band when the broadband average loudness approaches, matches, or exceeds the target. EE5. The method according to EE3, wherein step (a) includes determining an average loudness of each of the frequency bands, and step (b) includes applying the reduced DRC to each of the frequency bands whose average loudness approaches, matches, or exceeds the target. EE6. The method according to EE3, wherein determining the DRC gain includes smoothing the gain for each of the individual frequency bands across the individual frequency bands to improve timbre. A method of EE7, EE1, EE2, EE3, EE4, EE5, or EE6, wherein step (b) is: Determine the dynamic DRC gain gDRC; Smooth the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; Determine the dynamic gain g based on the minimum of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed: Including applying the dynamic gain g to the input audio signal, Method. A method according to EE8, wherein the dynamic gain g is: g = p * gDRC+(1 - p)*min(gDRC, gDRCsmoothed) Where "p" is a pumping parameter having a value in the range from 0 to 1. A method of EE1, EE2, EE3, EE4, EE5, EE6, EE7, or EE8, wherein the input audio signal has a loudness gradient, and the method further comprises: Controlling the release time constant for the application of reduced DRC and full DRC in response to the loudness gradient of the input audio signal. Method. The method according to EE9, wherein the release time constant is controlled to be faster in response to an increased steepness of the loudness gradient and slower in response to a decreased steepness of the loudness gradient. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, the method comprising: Determine an estimated level of the input audio signal; Determine the dynamic DRC gain gDRC by applying a DRC gain curve to the estimated level; Smooth the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; Determine the dynamic gain g based on the minimum determination of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed: Including applying the dynamic gain g to the input audio signal, thereby generating the output audio signal, Method. The method according to EE12.EE11, wherein the dynamic gain g is: g = p * gDRC+(1 - p)*min(gDRC, gDRCsmoothed) Where "p" is a pumping parameter having a value within the range from 0 to 1, Method. EE13. The input audio signal has a loudness gradient, and the method further comprises: Controlling the release time constant for applying DRC to the input audio signal in response to the loudness gradient of the input audio signal, The method according to EE11 or EE12. EE14. The method according to EE13, wherein the release time constant is controlled to be faster in response to an increased steepness of the loudness gradient and slower in response to a decreased steepness of the loudness gradient. EE15. The method according to EE13, wherein the control of the release time constant includes controlling the time constant for performing smoothing for generating the smoothed dynamic gain gDRCsmooothed. EE16. The method according to EE11, EE12, EE13, EE14 or EE15, wherein the input audio signal has a plurality of frequency bands, the dynamic gain g includes individual band gains for the individual ones of the frequency bands, and applying the dynamic gain g is: Including applying individual band gains to individual bands of the frequency bands of the input audio signal, Method. EE17. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: A level estimation subsystem coupled and configured to determine an estimated level of the input audio signal; coupled and configured to determine a dynamic DRC gain gDRC by applying a DRC gain curve to the level estimate value; coupled and configured to determine a dynamic gain g, including by smoothing the dynamic DRC gain gDRC to produce a smoothed dynamic gain gDRCsmoothed and determining the minimum value of each pair of corresponding values of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; and having a gain application subsystem coupled and configured to apply the dynamic gain g to an input audio signal to produce the output audio signal, wherein the gain determination subsystem determines the dynamic gain g such that when the dynamic gain g is applied to a segment of an input audio signal that includes a regular transient component, the system implements a first release time constant, and when the dynamic gain g is applied to a different segment of an input audio signal that does not include a regular transient component, the system implements a release time constant that is faster than the first release time constant. System. The system according to EE18.EE17, wherein the dynamic gain g is: g = p*gDRC+(1 - p)*min(gDRC,gDRCsmoothed) where "p" is a pumping parameter having a selectable value in the range from 0 to 1. EE19. The input audio signal has a loudness gradient, and the gain determination subsystem is configured to control a release time constant for application of DRC to the input audio signal in response to the loudness gradient of the input audio signal. The system according to EE17 or EE18. EE20. The system according to EE19, wherein the gain determination subsystem is configured to make the release time constant faster in response to an increased steepness of the loudness gradient and slower in response to a decreased steepness of the loudness gradient. EE21. The method according to EE19, wherein the gain determination subsystem is configured to control a release time constant, including by controlling a time constant for performing smoothing for generating the smoothed dynamic gain gDRCsmoothed. EE22. A system for performing dynamic range compression (DRC) on an input audio signal, comprising: a loudness determination subsystem coupled and configured to determine an average loudness of the input audio signal, wherein the average is over a time longer than a DRC application time of the DRC, and the DRC application time is an attack time or a release time of an instance of application of the DRC, or a duration of an instance of application of the DRC; a gain determination and application subsystem coupled and configured to apply a reduced DRC to the input audio signal when the average loudness of the input audio signal approaches, matches, or exceeds a target, thereby generating the output audio signal, or otherwise apply a full DRC to the input audio signal to generate the output audio signal. System. EE23. The system according to EE22, wherein the target is an audio signal level at least substantially equal to a knee point for the DRC or a maximum playback level of a playback system or device for playing back the output audio signal. EE24. The system according to EE22 or EE23, wherein the input audio signal has a plurality of frequency bands, and the gain determination and application subsystem is configured to determine a DRC gain for each of the individual frequency bands and apply the DRC gain to each of the individual frequency bands. EE25. The gain determination and application subsystem: determines a dynamic DRC gain gDRC; smooths the dynamic DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; Determine the dynamic gain g based on the minimum of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed: configured to apply the dynamic gain g to an input audio signal, a system according to EE22, EE23 or EE24. A system according to EE26.EE25, wherein the dynamic gain g is: g = p * gDRC+(1 - p)*min(gDRC, gDRCsmoothed) where "p" is a pumping parameter having a value in the range from 0 to 1. A system according to EE27.EE22, EE23, EE24, EE25, or EE26, wherein the input audio signal has a loudness gradient and the gain determination and application subsystem is: configured to control a release time constant for application of reduced DRC and full DRC in response to the loudness gradient of the input audio signal. System. EE28. The system according to EE22, EE23, EE24, EE25, EE26 or EE27, wherein the gain determination and application subsystem is configured to make the release time constant faster in response to an increased steepness of the loudness gradient and slower in response to a decreased steepness of the loudness gradient. Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art. The general principles defined herein may be applied to other implementations without departing from the spirit or scope of the disclosure. Accordingly, the claims are not intended to be limited to the specific implementations described and shown herein, but should be accorded the widest scope consistent with the disclosure.

[0077] The methods and systems described in this disclosure may be implemented as software, firmware, and / or hardware. For example, certain components (e.g., each of elements 1, 2, 4, 6, and 7 of FIG. 2, or each of elements 1, 4, 6, 7, and 8 of FIG. 2A, or each of elements 1, 3, 5, 7, 11, and 13 of FIG. 3, or each of elements 1, 3, 5, 7, 11, 13, and 15 of FIG. 4) may be implemented as software operating on a digital signal processor (e.g., having an input coupled to receive an input audio signal) or a microprocessor. Some components may be implemented as hardware and / or as application specific integrated circuits. Signals encountered in the methods and systems described above may be stored in a medium such as random access memory or an optical storage medium. They may be transferred via a network such as a radio network, a satellite network, a wireless network, or a wired network such as the Internet. A typical apparatus utilizing the methods and systems described in this disclosure is a portable electronic device or other consumer equipment used to store and process (e.g., implement playback or rendering) an audio signal (e.g., an output audio signal generated according to any embodiment of the system or method of the present invention). In this application, the description in the form of "A and / or B" means "A" or "B" or "A and B".

Claims

1. A method for performing dynamic range compression (DRC) on an input audio signal to generate an output audio signal, comprising: (a) determining an average level or power of the input audio signal, wherein the average is over a time longer than an attack time and / or a release time of the application of the DRC that applies a gain other than 1; (b) generating the output audio signal without applying DRC to the input audio signal when the average level or power of the input audio signal approaches, matches, or exceeds a target, or otherwise applying the DRC to the input audio signal to generate the output audio signal, wherein the DRC increases the level or power of the input audio signal when the level or power of the input audio signal is below a knee-point level or power, and the target is an audio signal level equal to the knee-point for the DRC or the maximum playback level of a playback system or device for playing back the output audio signal; Applying the DRC to the input audio signal in step (b) to generate the output audio signal comprises: determining a DRC gain gDRC; smoothing the DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; determining a dynamic gain g based on the smaller of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed; applying the dynamic gain g to the input audio signal. Method.

2. The method according to claim 1, wherein the input audio signal has a plurality of frequency bands.

3. The method according to claim 2, wherein step (b) comprises determining a DRC gain for at least one of the plurality of frequency bands and applying the DRC gain to the frequency band.

4. The method according to claim 2, wherein step (b) comprises determining a DRC gain for each of the individual frequency bands and applying the DRC gain determined for each of the individual frequency bands.

5. Step (a) includes determining the broadband average level or power of the input audio signal, and step (b) includes not applying DRC to each frequency band when the broadband average level or power approaches or equals or exceeds the target, the method according to claim 2.

6. Step (a) includes determining the respective average level or power for each of the frequency bands, and step (b) includes not applying DRC to each frequency band where the average level or power approaches or equals or exceeds the target, the method according to claim 2.

7. The method according to any one of claims 1 to 6, wherein the dynamic gain g is: g = p * gDRC + (1 - p) * min(gDRC, gDRCsmoothed) where "p" is a pumping parameter having a value within the range from 0 to 1.

8. The method according to claim 7, wherein the value of the pumping parameter within the range from 0 to 1 is selectable by a user.

9. A system for performing dynamic range compression DRC on an input audio signal, comprising: A loudness determination subsystem coupled and configured to determine the average level or power of the input audio signal, wherein the average is over a time longer than the attack time and / or release time of the application of the DRC that applies a gain other than 1; A gain determination and application subsystem coupled and configured to generate an output audio signal without applying DRC to the input audio signal when the average level or power of the input audio signal approaches or equals or exceeds a target, and otherwise to apply the DRC to the input audio signal to generate the output audio signal, wherein the DRC raises the level or power of the input audio signal when the level or power of the input audio signal is below the knee-point level or power, and the target is an audio signal level equal to the knee-point for the DRC or the maximum playback level of a playback system or device for playing back the output audio signal. The gain determination and application subsystem configured to generate the output audio signal by applying the DRC to the input audio signal: Determine the DRC gain gDRC; Smooth the DRC gain gDRC to generate a smoothed dynamic gain gDRCsmoothed; Determine a dynamic gain g based on the smaller of the DRC gain gDRC and the smoothed dynamic gain gDRCsmoothed: Including applying the dynamic gain g to the input audio signal, System.

10. The system according to claim 9, wherein the input audio signal has a plurality of frequency bands.

11. The system according to claim 10, wherein the gain determination and application subsystem is configured to determine a DRC gain for at least one of the plurality of frequency bands and apply the DRC gain to the frequency band.

12. The system according to claim 10, wherein the gain determination and application subsystem is configured to determine a DRC gain for each individual one of the frequency bands and apply the DRC gain determined for each individual one of the frequency bands.

13. The system according to claim 10, wherein the gain determination and application subsystem determines a broadband average level or power of the input audio signal and does not apply DRC to each frequency band when the broadband average level or power approaches or matches or exceeds the target.

14. The system according to claim 10, wherein the gain determination and application subsystem determines an average level or power for each of the frequency bands and does not apply DRC to each frequency band whose average level or power approaches or matches or exceeds the target.

15. The dynamic gain g is: g = p * gDRC + (1 - p) * min(gDRC, gDRCsmoothed) where "p" is a pumping parameter having a value in the range from 0 to 1, The system according to any one of claims 9 to 14.

Citation Information

Patent Citations

  • Calculation and adjustment of perceived loudness and / or perceived spectral balance of audio signals

    JP2008518565A

  • Audio compressor

    JP2012065068A

  • Method for compressing the dynamics in an audio signal

    US20160381468A1

  • Dynamic range control apparatus

    WO2013038451A1