Dynamic range control for diverse playback environments

The audio encoder's dynamic range compression and differential gain approach allows decoders to customize audio processing for diverse playback environments, ensuring consistent loudness and intelligibility, addressing the challenge of varying playback conditions.

JP2026021423APending Publication Date: 2026-02-10DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025182205
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-02-10
Filing Date
2025-10-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Media processing devices struggle to maintain consistent loudness and intelligibility across diverse media formats and content types, particularly when playing high-quality, wide-bandwidth, and wide-dynamic-range audio content in various playback environments.

Method used

An audio encoder transmits encoded audio signals with dynamic range compression curves and differential gains, allowing audio decoders to customize audio processing based on the playback environment, using techniques like auditory scene analysis and gain profiling to ensure consistent loudness and intelligibility.

Benefits of technology

The solution enables flexible and efficient dynamic range control, maintaining perceptual quality and consistent loudness across different playback environments, supporting various playback devices and configurations without requiring the decoder to perform complex gain calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021423000001_ABST
    Figure 2026021423000001_ABST
Patent Text Reader

Abstract

To provide a method, a storage medium, an apparatus and a program for providing dynamic range control (DRC) for various reproduction environments.SOLUTION: An audio decoder (100) operating in a playback channel configuration different from a reference channel configuration, wherein an encoded input signal (102) comprises audio sample data and DRC metadata for each reference channel, the DRC metadata comprising a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration. The audio decoder 100 selects, from the DRC gains for the plurality of channel configurations, a set of DRC gains for a playback channel configuration to apply as part of an overall gain applied to the audio sample data to generate output audio sample data for each channel of the playback channel configuration. The reproduction channel configuration is a two channel configuration.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 61 / 877,230, filed September 12, 2013, U.S. Provisional Patent Application No. 61 / 891,324, filed October 15, 2013, and U.S. Provisional Patent Application No. 61 / 938,043, filed February 10, 2014, the contents of each of which are incorporated herein by reference in their entirety.

[0002] technology The present invention relates generally to processing audio signals, and more particularly to techniques that can be used to apply dynamic range control and other types of audio processing operations to audio signals in any of a wide variety of playback environments. [Background technology]

[0003] The growing popularity of media consumption devices has created new opportunities and challenges for creators and distributors of media content for playback on such devices, and for designers and manufacturers of such devices. Many consumer devices can play a wide range of media content types and formats, including those often associated with high-quality, wide-bandwidth, and wide-dynamic-range audio content for HDTV, Blu-ray, or DVD. Media processing devices can be used to play this type of audio content over their own internal acoustic transducers or over external transducers such as headphones. However, media processing devices generally cannot play this content with consistent loudness and intelligibility across diverse media formats and content types.

[0004] The approaches described in this section could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, unless otherwise noted, it should not be assumed that problems identified with one or more approaches have been recognized by the prior art based on this section. [Brief explanation of the drawings]

[0005] The present invention is illustrated by way of example, and not by way of limitation, in the accompanying figures in which like reference symbols refer to similar elements and in which: [Figure 1A] FIG. 1 illustrates an exemplary audio decoder. [Figure 1B] FIG. 1 illustrates an exemplary audio encoder. [Figure 2A] FIG. 2 illustrates an exemplary dynamic range compression curve. [Figure 2B] FIG. 2 illustrates an exemplary dynamic range compression curve. [Figure 3] FIG. 10 illustrates exemplary processing logic for combined DRC and limiting gain determination / calculation. [Figure 4] FIG. 10 illustrates an exemplary differential encoding of gain. [Figure 5] FIG. 1 illustrates an exemplary codec system having an audio encoder and an audio decoder. [Figure 5A] FIG. 1 illustrates an exemplary audio decoder. [Figure 6A] FIG. 1 illustrates an exemplary process flow. [Figure 6B] FIG. 1 illustrates an exemplary process flow. [Figure 6C] FIG. 1 illustrates an exemplary process flow. [Figure 6D] FIG. 1 illustrates an exemplary process flow. [Figure 7] FIG. 1 illustrates an exemplary hardware platform upon which a computer or computing device described herein may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0006] Illustrative embodiments relating to applying dynamic range control and other types of audio processing operations to audio signals in any of a wide variety of playback environments are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without such specific details. Conversely, well-known structures and devices are not described in exhaustive detail to avoid unnecessarily obscuring, obscuring, or obscuring the present invention.

[0007] Exemplary embodiments are described herein according to the following outline. 1. General Overview 2. Dynamic range control 3. Audio Decoder 4. Audio Encoder 5. Dynamic Range Compression Curve 6. DRC Gain, Gain Limiting and Gain Smoothing 7. Input and Gain Smoothing 8. DRC across multiple frequency bands 9. Volume adjustment in the loudness area 10. Gain profile with differential gain 11. Additional Gain-Related Actions 12. Specific and broadband loudness levels 13. Individual gains for individual subsets of channels 14. Auditory Scene Analysis 15. Loudness Level Transition 16. Reset 17. Gain provided by the encoder 18. Exemplary System and Process Flows 19. Implementation Mechanism - Hardware Overview 20. Equivalents, Extensions, Substitutions, etc.

[0008] 1. General Overview This overview provides a basic description of some aspects of embodiments of the present invention. It should be noted that this overview is not a comprehensive or exhaustive summary of aspects of the embodiments. Furthermore, it should be noted that this overview is not intended to identify any particularly significant aspects or elements of the embodiments, nor to delineate any scope of the invention generally, and in particular of the embodiments. This overview merely presents some concepts related to the exemplary embodiments in a condensed and simplified form and should merely be understood as a conceptual prelude to the more detailed description of the exemplary embodiments that follows. It should be noted that although separate embodiments are discussed herein, any combination of the embodiments and / or sub-embodiments discussed herein may be combined to form further embodiments.

[0009] In some approaches, an encoder assumes that the audio content is being encoded for a particular environment for the purposes of dynamic range control and determines audio processing parameters, such as gain, for that particular environment. The gain determined by the encoder under these approaches is typically smoothed over some time interval, with some time constant (e.g., an exponential decay function). Furthermore, the gain determined by the encoder under these approaches may incorporate gain limiting to ensure that the signal does not exceed the clipping level for the assumed environment. Thus, the gain encoded by the encoder into the audio signal along with the audio information under these approaches is the result of many different effects and is irreversible. A decoder receiving the gain under these approaches would be unable to distinguish which portion of the gain is for dynamic range control, which portion of the gain is for gain smoothing, and which portion of the gain is for gain limiting.

[0010] Under the techniques described herein, an audio encoder does not assume that a particular playback environment at an audio decoder is supported. In one embodiment, the audio encoder transmits an encoded audio signal having audio content from which a correct loudness level (e.g., without clipping) can be determined. The audio encoder may also transmit one or more dynamic range compression curves to the audio decoder. Any of the one or more dynamic range compression curves may be standards-based, proprietary, customized, content provider-specific, etc. Reference loudness levels, attack times, release times, etc. may be transmitted by the audio encoder as part of or in association with the one or more dynamic range compression curves.

[0011] In some embodiments, the audio encoder implements an auditory scene analysis (ASA) technique, uses the ASA technique to detect auditory events in the audio content, and sends one or more ASA parameters describing the detected auditory events to the audio decoder.

[0012] In some embodiments, the audio encoder may also be configured to detect reset events in the audio content and send an indication of the reset event together with the audio content in a time-synchronous manner to a downstream device, such as an audio decoder.

[0013] In some embodiments, an audio encoder may be configured to calculate one or more sets of gains (e.g., DRC gains) for individual portions of audio content (e.g., audio data blocks, audio data frames, etc.) and encode the sets of gains along with the individual portions of audio content into an encoded audio signal. In some embodiments, the sets of gains generated by the audio encoder correspond to one or more gain profiles (e.g., as shown in Table 1). In some embodiments, Huffman coding, differential coding, etc. may be used to encode or retrieve the sets of gains into or from components, subdivisions, etc. of audio data frames. These components, subdivisions, etc. may be referred to as subframes of an audio data frame. Different sets of gains may correspond to different sets of subframes. Each set of gains or each set of subframes may have two or more temporal components (e.g., subframes, etc.). In some embodiments, a bitstream formatter in an audio encoder described herein may use one or more for loops to write one or more sets of gains together as differential data codes into one or more sets of subframes in an audio data frame. Correspondingly, a bitstream parser in an audio decoder described herein may read any of the one or more sets of gains encoded as the differential data codes from the one or more sets of subframes in an audio data frame.

[0014] In some embodiments, the audio encoder determines a dialogue loudness level in the audio content to be encoded into the encoded audio signal and sends the dialogue loudness level along with the audio content to the audio decoder.

[0015] In some embodiments, an audio encoder sends to a downstream receiving audio decoder a default dynamic compression curve for a default gain profile in a playback environment or scenario. In some embodiments, an audio encoder assumes that the downstream receiving audio decoder uses the default dynamic compression curve for the default gain profile in a playback environment or scenario. In some embodiments, an audio encoder sends an indication to a downstream receiving audio decoder as to which of one or more dynamic compression curves defined in the downstream receiving audio decoder should be used in a playback environment or scenario. In some embodiments, for each of one or more non-default gain profiles, the audio encoder sends a (e.g., non-default) dynamic compression curve corresponding to that non-default profile as part of the metadata carried by the encoded audio signal. The techniques described herein allow multiple sets of differential gains relative to a default compression curve to be generated by an upstream encoder and sent to a downstream decoder. This allows great freedom in the design of the DRC compressor in the decoder (e.g., the process for calculating gains based on the compression curves and the smoothing operation, etc.), while keeping the required bitrate relatively low compared to transmitting full gain values. For purposes of illustration only, a default profile or default DRC curve is referred to as one relative to which differential gains for non-default profiles or non-default DRC curves can be specifically calculated. However, this is merely for purposes of illustration, and there is no strict need to distinguish between default and non-default profiles (e.g., in a media data stream), as in various embodiments, all other profiles may be differential gains relative to the same particular (e.g., "default") compression curve. As used herein, a "gain profile" may also refer to a DRC mode as an operating mode of a compressor that performs a DRC operation.In some embodiments, a DRC mode relates to a specific type of playback device (AVR, TV, tablet) and / or environment (noisy, quiet, late night). Each DRC mode can be associated with a gain profile. The gain profile may be represented by definition data based on which the compressor performs the DRC operation. In some embodiments, the gain profile can be a DRC curve (possibly parameterized) and time constants used in the DRC operation. In some embodiments, the gain profile can be a collection of DRC gains as the output of the DRC operation in response to the audio signal. The profiles of different DRC modes may correspond to different amounts of compression.

[0016] In some embodiments, the audio encoder determines a set of default (e.g., full DRC and no DRC, full DRC, etc.) gains for audio content based on a default dynamic range compression curve corresponding to a default gain profile, and determines a set of non-default (e.g., full DRC and no DRC, full DRC, etc.) gains for the same audio content for each of one or more non-default gain profiles. The audio encoder may then determine a gain difference between the set of default (e.g., full DRC and no DRC, full DRC, etc.) gains for the default gain profile and the set of non-default (e.g., full DRC and no DRC, full DRC, etc.) gains for the non-default gain profile, include the gain difference in a set of differential gains, etc. Instead of sending a (e.g., non-default) dynamic range compression curve for a non-default profile associated with a non-default playback environment or scenario, the audio encoder may send the set of differential gains as part of the metadata carried by the encoded audio signal instead of or in addition to the non-default dynamic range compression curve.

[0017] The set of differential gains may be smaller in size than the set of non-default (e.g., full DRC and no DRC, full DRC, etc.) gains. In this manner, transmitting differential gains rather than non-differential (e.g., full DRC and no DRC, full DRC, etc.) gains may require a lower bit rate than transmitting the non-differential (e.g., full DRC and no DRC, full DRC, etc.) gains directly.

[0018] Audio decoders that receive the encoded audio signals described herein may be provided by different manufacturers and implemented with different components and designs. The audio decoders may have been released to end users at different times or updated with different versions of hardware, software, or firmware. As a result, the audio decoders may have different audio processing capabilities. In some embodiments, multiple audio decoders may be capable of supporting a limited set of gain profiles, such as a default gain profile defined by a standard, proprietary requirements, etc. Multiple audio decoders may be configured with the capability to perform associated gain generation operations to generate gain for a default gain profile based on a default dynamic range compression curve representing the default gain profile. Transmitting a default dynamic range compression curve for a default gain profile in an audio signal may be more efficient than transmitting a generated / calculated gain for the default gain profile in an audio signal.

[0019] On the other hand, for a non-default gain profile, the audio encoder can pre-generate a differential gain by referencing a particular default dynamic range compression curve corresponding to the particular default gain profile. In response to receiving the differential gain in an audio signal generated by the audio encoder, the audio decoder can render the received audio content by generating a default gain based on the default dynamic range compression curve received in the audio signal, combining the received differential gain and the generated default gain into a non-default gain for the non-default gain profile, applying the non-default gain to audio content decoded from the audio signal, etc. In some embodiments, the non-default gain profile can be used to ensure the constraints of the default dynamic range compression curve.

[0020] The techniques described herein can be used to provide flexible support for new gain profiles, features, or enhancements. In some embodiments, at least one gain profile cannot be easily represented using a default or non-default dynamic range compression curve. In some embodiments, at least one gain profile may be specific to particular audio content (e.g., a particular movie). It may also be that a representation of a non-default gain profile (e.g., a parameterized DRC curve) requires transmitting more parameters, smoothing constants, etc. in the encoded audio signal than can be carried in the encoded audio signal. In some embodiments, at least one gain profile may be specific to a particular audio content provider (e.g., a particular studio).

[0021] In this way, the audio encoder described herein can take the initiative in supporting a new gain profile by implementing gain generation operations for the new gain profile and for the default gain profile to which the new gain profile relates. A downstream receiving audio decoder does not need to perform gain generation operations for the new gain profile. Rather, the audio decoder can support the new gain profile by leveraging the non-default differential gain generated by the audio encoder without the audio decoder performing gain generation operations for the new gain profile.

[0022] In some embodiments, profile-related metadata encoded in the encoded audio signal includes one or more (e.g., default, etc.) dynamic range compression curves and one or more sets of (e.g., non-default, etc.) differential gains, structured, indexed, etc., according to respective gain profiles to which the one or more (e.g., default, etc.) dynamic range compression curves and one or more sets of (e.g., non-default, etc.) differential gains correspond. In some embodiments, a relationship between a set of non-default differential gains and a default dynamic range compression curve may be indicated in the profile-related metadata. This may be particularly useful when more than one default dynamic range compression curve is present in the metadata, or is not present in the metadata but is defined in a downstream decoder. Based on the relationship indicated in the profile-related metadata, a receiving audio decoder can determine which default dynamic range compression curve should be used to generate a set of default gains. The generated gains can then be combined with the received set of non-default differential gains to generate non-default gains, e.g., to compensate for limitations of the default dynamic range compression curves.

[0023] The techniques described herein do not require the audio decoder to be locked in with any (e.g., lossy) audio processing that may have been performed by an upstream device such as an audio encoder, while assuming a hypothetical playback environment, scenario, etc. in a hypothetical audio decoder. The decoders described herein may be configured to customize audio processing operations based on the particular playback scenario, for example, to distinguish between various loudness levels present in the audio content, minimize loss in audio perceptual quality at or near boundary loudness levels, maintain spatial balance among channels or subsets of channels, etc.

[0024] An audio decoder receiving an encoded audio signal with dynamic range compression curves, reference loudness levels, attack times, release times, etc. can determine the particular playback environment being used at the decoder and select a particular compression curve with a corresponding reference loudness level that corresponds to that particular playback environment.

[0025] The decoder can calculate / determine the loudness level of each portion of the audio content (e.g., audio data block, audio data frame, etc.) extracted from the encoded audio signal, or obtain the loudness level of each portion of the audio content if the audio encoder calculated and provided the loudness level in the encoded audio signal. Based on one or more of the loudness level of each portion of the audio content, the loudness level of previous portions of the audio content, the loudness level of subsequent portions of the audio content if available, the specific compression curve, a specific profile related to the specific playback environment or scenario, etc., the decoder determines audio processing parameters such as gain for dynamic range control (DRC gain), attack time, release time, etc. The audio processing parameters can also include adjustments to align the dialogue loudness level to a specific reference loudness level for the specific playback environment (which may be user-adjustable).

[0026] The decoder uses the audio processing parameters to apply audio processing operations including dynamic range control (e.g., multi-channel, multi-band, etc.), dialogue level adjustment, etc. Audio processing operations performed by the decoder may further include, but are not limited to, gain smoothing based on attack and release times provided as part of or in conjunction with a selected dynamic range compression curve, gain limiting to prevent clipping, etc. Different audio processing operations may be performed with different (e.g., adjustable, threshold-dependent, controllable, etc.) time constants. For example, gain limiting to prevent clipping may be applied to individual audio data blocks, individual audio data frames, etc. with a relatively short time constant (e.g., instantaneous, approximately 5.3 milliseconds, etc.).

[0027] In some embodiments, the decoder may be configured to extract ASA parameters (e.g., temporal positions of auditory event boundaries, time-dependent values ​​of event certainty indicators, etc.) from metadata in the encoded audio signal and to control the speed of gain smoothing at auditory events (e.g., using short time constants for attack at auditory event boundaries, using long time constants to slow down gain smoothing within auditory events, etc.) based on the extracted ASA parameters.

[0028] In some embodiments, the decoder also maintains a histogram of instantaneous loudness levels over a time interval or window and uses the histogram to control the rate of gain change at loudness level transitions such as between programs, between programs and commercials, for example by modifying the time constant.

[0029] In some embodiments, the decoder supports two or more speaker configurations (e.g., a portable mode with speakers, a portable mode with headphones, a stereo mode, a multi-channel mode, etc.). The decoder may be configured, for example, to maintain the same loudness level between two different speaker configurations (e.g., between stereo and multi-channel modes) when playing the same audio content. The audio decoder may use one or more downmixing equations to downmix multi-channel audio content received from an audio signal encoded for a reference speaker configuration to a particular speaker configuration in the audio decoder, the multi-channel audio content being encoded for the reference speaker configuration.

[0030] In some embodiments, automatic gain control (AGC) may be disabled in the audio decoders described herein.

[0031] In some embodiments, the device may be part of a media processing system, including, but not limited to, audiovisual devices, flat panel TVs, handheld devices, game consoles, televisions, home theater systems, tablets, mobile devices, laptop computers, netbook computers, cellular wireless telephones, e-readers, point-of-sale terminals, desktop computers, computer workstations, computer kiosks, and various other types of terminals and media processing units.

[0032] Various modifications to the preferred embodiment and general principles and features described herein will be readily apparent to those skilled in the art, and thus the present disclosure is not intended to be limited to the embodiments shown but is to be accorded the widest scope consistent with the principles and features described herein.

[0033] 2. Dynamic range control Without customized dynamic range control, the input audio information (e.g., PCM samples, time-frequency samples in a QMF matrix, etc.) is often reproduced at a loudness level that is inappropriate for the playback device's particular playback environment (i.e., including the device's physical and / or mechanical playback limitations), which may differ from the playback environment that was targeted when the encoded audio content was coded in the encoding device.

[0034] The techniques described in this paper can be used to support dynamic range control of a wide variety of audio content customized for any of a wide variety of playback environments, while maintaining the perceptual quality of the audio content.

[0035] Dynamic range control (DRC) refers to a time-dependent audio processing operation that changes (e.g., compresses, cuts, expands, boosts, etc.) the input dynamic range of loudness levels in audio content into an output dynamic range that differs from the input dynamic range. For example, in a dynamic range control scenario, soft sounds may be mapped (e.g., boosted) to higher loudness levels, and loud sounds may be mapped (e.g., cut) to lower loudness values. As a result, in the loudness domain, in this example, the output range of loudness levels is smaller than the input range of loudness levels. However, in some embodiments, dynamic range control may be reversible so that the original range is restored. For example, an expansion operation may be performed to restore the original range, as long as the mapped loudness levels in the output dynamic range mapped from the original loudness levels are below the clipping level, each unique original loudness level is mapped to a unique output loudness level, etc.

[0036] The DRC techniques described herein can be used to provide a better listening experience in certain playback environments or situations. For example, soft sounds in a noisy environment may be masked by noise that makes the soft sounds inaudible. Conversely, loud sounds may be undesirable in some situations, such as with noisy neighbors. Many devices, typically with small form factor loudspeakers, cannot reproduce sound at high output levels. In some cases, lower signal levels may be reproduced below the human hearing threshold. The DRC techniques may map input loudness levels to output loudness levels based on DRC gains (e.g., scaling factors for scaling audio amplitude, boost ratios, cut ratios, etc.) found using a dynamic range compression curve.

[0037] A dynamic range compression curve is a function (e.g., a look-up table, a curve, a multi-segment piecewise linear function, etc.) that maps individual input loudness levels (e.g., non-dialogue sounds) determined from individual audio data frames to individual gains or gains for dynamic range control. Each individual gain indicates the amount of gain to be applied to the corresponding individual input loudness level. The output loudness level after applying the individual gains represents the target loudness level for the audio content in that individual audio data frame in a specific playback environment.

[0038] In addition to specifying a mapping between gain and loudness level, a dynamic range compression curve may include or be provided with specific release and attack times when applying a specific gain. Attack refers to an increase in signal energy (or loudness) between successive time samples, while release refers to a decrease in signal energy (or loudness) between successive time samples. The attack time (e.g., 10 ms, 20 ms, etc.) refers to the time constant used to smooth the DRC gain when the corresponding signal is in attack mode. The release time (e.g., 80 ms, 100 ms, etc.) refers to the time constant used to smooth the DRC gain when the corresponding signal is in release mode. In some embodiments, these time constants are additionally, optionally, or alternatively used to smooth the signal energy (loudness) before determining the DRC gain.

[0039] Different playback environments may correspond to different dynamic range compression curves. For example, a dynamic range compression curve for a playback environment of a flat-panel TV may be different from a dynamic range compression curve for a playback environment of a portable device. In some embodiments, a playback device may have more than one playback environment. For example, a first dynamic range compression curve for a first playback environment of a portable device using speakers may be different from a second dynamic range compression curve for a second playback environment of the same portable device using a headset.

[0040] 3. Audio Decoder FIG. 1A shows an exemplary audio decoder 100 having a data extractor 104, a dynamic range controller 106, an audio renderer 108, and the like.

[0041] In some embodiments, the data extractor (104) is configured to receive the encoded input signal 102. The encoded input signal described herein may be a bitstream containing encoded (e.g., compressed) input audio data frames and metadata. The data extractor (104) is configured to extract / decode the input audio data frames and metadata from the encoded input signal (102). Each input audio data frame has multiple coded audio data blocks, each representing multiple audio samples. Each frame represents a (e.g., fixed) time interval containing a certain number of audio samples. The frame size may vary with the sample rate and the coding data rate. An audio sample is a quantized audio data element (e.g., an input PCM sample, an input time-frequency sample in a QMF matrix, etc.) that represents spectral content in one, two, or more (audio) frequency bands or frequency ranges. The quantized audio data elements in the input audio data frame may represent pressure waves in the digital (quantized) domain. Quantized audio data elements may cover a finite range of loudness levels up to a maximum possible value (e.g., clipping level, maximum loudness level, etc.).

[0042] The metadata can be used by a wide variety of receiving decoders to process the input audio data frames. The metadata may include various operational parameters related to one or more operations to be performed by the decoder (100), normalization parameters related to the dialogue loudness level represented in the input audio data frames, etc. Dialogue loudness level may refer to the dialogue loudness, program loudness, average dialogue loudness, etc. (e.g., psychoacoustic, perceptual, etc.) level of an entire program (e.g., a movie, television program, radio broadcast, etc.), a portion of a program, the dialogue of a program, etc.

[0043] The operation and functionality of the decoder (100) or some or all of its modules (e.g., data extractor 104, dynamic range controller 106, etc.) may be adapted in response to metadata extracted from the encoded input signal (102). For example, the metadata—including, but not limited to, dynamic range compression curves, dialogue loudness levels, etc.—may be used by the decoder (100) to generate output audio data elements in the digital domain (e.g., output PCM samples, output time-frequency samples in a QMF matrix, etc.). The output data elements can then be used to drive audio channels or speakers to achieve a specified loudness or reference playback level during playback in a particular playback environment.

[0044] In some embodiments, the dynamic range controller (106) is configured to receive some or all of the audio data elements and metadata in the input audio data frames, and to perform audio processing operations (e.g., dynamic range control operations, gain smoothing operations, gain limiting operations, etc.) on the audio data elements in the input audio data frames based at least in part on the metadata extracted from the encoded audio signal (102).

[0045] In some embodiments, the dynamic range controller (106) may include a selector 110, a loudness calculator 112, a DRC gain unit 114, etc. The selector (110) may be configured to determine a speaker configuration associated with a particular playback environment in the decoder (100) (e.g., flat panel mode, portable device with speakers, portable device with headphones, 5.1 speaker configuration, 7.1 speaker configuration, etc.), select a particular dynamic range compression curve from the dynamic range compression curves extracted from the encoded input signal (102), etc.

[0046] The loudness calculator (112) may be configured to calculate one or more types of loudness levels represented by audio data elements in the input audio data frame. Exemplary types of loudness levels include, but are not limited to, any of the following: individual loudness levels across individual frequency bands in individual channels across individual time intervals; broadband loudness levels across a wide frequency range in individual channels; loudness levels determined from or smoothed across an audio data block or frame; loudness levels determined from or smoothed across two or more audio data blocks or frames; and loudness levels smoothed across one or more time intervals. Zero, one, or more of these loudness levels may be modified by the decoder (100) for dynamic range control.

[0047] To determine the loudness level, the loudness calculator (112) can determine one or more time-dependent physical wave attributes, such as spatial pressure levels at specific audio frequencies, represented by audio data elements in the input audio data frame. The loudness calculator (112) can use the one or more time-varying physical wave attributes to derive one or more types of loudness levels based on one or more psychoacoustic functions that model human loudness perception. The psychoacoustic functions can be nonlinear functions, such as those constructed based on a model of the human auditory system, that convert specific spatial pressure levels at specific audio frequencies into specific loudness for the specific audio frequencies.

[0048] A loudness level across multiple (audio frequencies) or multiple frequency bands (e.g., broadband, wideband, etc.) may be derived through integration of specific loudness levels across multiple (audio) frequencies or multiple frequency bands. A time-averaged, smoothed, etc. loudness level across one or more time intervals (e.g., longer than represented by the audio data elements in an audio data block or frame) may be obtained using one or more smoothing filters implemented as part of the audio processing operations in the decoder (100).

[0049] In an exemplary embodiment, specific loudness levels for different frequency bands may be calculated for each audio data block of a certain number (e.g., 256 samples). A pre-filter may be used to apply frequency weighting (e.g., similar to IEC B weighting) to the specific loudness levels before integrating them into a broadband loudness level. A summation of broad loudness levels across two or more channels (e.g., left front, right front, center, left surround, right surround, etc.) may be performed to provide an overall loudness level for the two or more channels.

[0050] In some embodiments, an overall loudness level may refer to a broadband loudness level in a channel (e.g., center) of a speaker configuration. In some embodiments, an overall loudness level may refer to a broadband loudness level in multiple channels. The multiple channels may be all channels in a speaker configuration. Additionally, optionally, or alternatively, the multiple channels may include a subset of channels in a speaker configuration (e.g., a subset of channels including left front, right front, and low frequency effects (LFE), a subset of channels including left surround and right surround, a subset of channels including center, etc.).

[0051] The loudness levels (e.g., broadband, wideband, global, specific, etc.) may be used as inputs to find corresponding DRC gains (e.g., static, pre-smoothing, pre-limiting, etc.) from a selected dynamic range compression curve. The loudness levels used as inputs to find the DRC gains may first be adjusted or normalized with respect to the dialogue loudness levels from metadata extracted from the encoded audio signal (102). In some embodiments, the adjustment and normalization related to the adjustment of the dialogue loudness levels may be performed in a non-loudness domain (e.g., SPL domain) for a portion of the audio content in the encoded audio signal (102), before a specific spatial pressure level represented in said portion of the audio content in the encoded audio signal (102) is converted or mapped to a specific loudness level for said portion of the audio content in the encoded audio signal (102).

[0052] In some embodiments, the DRC gain unit (114) may be configured with a DRC algorithm to generate a gain (e.g., for dynamic range control, gain limiting, gain smoothing, etc.) and apply the gain to one or more loudness levels at one or more types of loudness levels represented by audio data elements in the input audio data frame to achieve a target loudness level for the particular playback environment. The application of a gain (e.g., a DRC gain, etc.) as described herein may, but need not, occur in the loudness domain. In some embodiments, a gain may be generated based on a loudness calculation (which may be expressed in terms of SPL, or simply compensated for, e.g., dialogue loudness levels without conversion), smoothed, and applied directly to the input signal. In some embodiments, a technique as described herein may apply a gain to a signal in the loudness domain, then convert the signal from the loudness domain back to the (linear) SPL domain, and calculate a corresponding gain to be applied to the signal by evaluating the signal in the loudness domain before and after the gain has been applied to the signal. The ratio (or difference when expressed in log-dB notation) then determines the corresponding gain for that signal.

[0053] In some embodiments, the DRC algorithm operates in conjunction with multiple DRC parameters. The DRC parameters include a dialogue loudness level that has already been calculated and embedded in the encoded audio signal (102) by an upstream encoder (e.g., 150) and can be obtained by the decoder (100) from metadata in the encoded audio signal (102). The dialogue loudness level from the upstream encoder indicates an average dialogue loudness level (e.g., per program, relative to the energy of a full-scale 1 kHz sine wave, relative to the energy of a reference square wave, etc.). In some embodiments, the dialogue loudness level extracted from the encoded audio signal (102) may be used to reduce loudness level differences between programs. In one embodiment, the reference dialogue loudness level may be set to the same value for different programs in the same specific playback environment in the decoder (100). Based on the dialogue loudness level from the metadata, the DRC gain unit (114) can apply a dialogue loudness-related gain to each audio data block in the program such that the output dialogue loudness level averaged across multiple audio data blocks of the program is raised / lowered to a reference dialogue loudness level (e.g., pre-configured, system default, user-configurable, profile-dependent, etc.) for that program.

[0054] In some embodiments, DRC gains may be used to address differences in loudness levels within a program by boosting or cutting signal portions that are soft and / or loud according to a selected dynamic range compression curve. One or more of these DRC gains may be calculated / determined by a DRC algorithm based on the selected dynamic range compression curve and the loudness level (broadband, wideband, overall, specific, etc.) determined from one or more corresponding audio data blocks, audio data frames, etc.

[0055] The loudness level used to determine the DRC gain (e.g., static, pre-smoothing, pre-gain-limiting, etc.) by searching a selected dynamic range compression curve may be calculated over a short interval (e.g., approximately 5.3 milliseconds). The integration time of the human auditory system can be much longer (e.g., approximately 200 milliseconds). The DRC gain obtained from the selected dynamic range compression curve may be smoothed with a time constant to account for the long integration time of the human auditory system. To implement a fast rate of change (increase or decrease) in loudness level, a short time constant may be used to cause the loudness level to change over a short time interval corresponding to the short time constant. Conversely, to implement a slow rate of change (increase or decrease) in loudness level, a long time constant may be used to change the loudness level over a long time interval corresponding to the long time constant.

[0056] The human auditory system may respond to increasing and decreasing loudness levels with different integration times. In some embodiments, different time constants may be used to smooth the static DRC gain retrieved from the selected dynamic range compression curve, depending on whether the loudness level is increasing or decreasing. For example, to accommodate the characteristics of the human auditory system, attacks (increases in loudness level) may be smoothed with a relatively short time constant (e.g., attack time), while releases (decrease in loudness level) may be smoothed with a relatively long time constant (e.g., release time).

[0057] A DRC gain for a portion of audio content (e.g., one or more audio data blocks, audio data frames, etc.) may be calculated using a loudness level determined from the portion of audio content. The loudness level to be used for lookup in the selected dynamic range compression curve may first be adjusted with respect to (e.g., relative to) a dialogue loudness level (e.g., of a program of which the audio content is a part) in metadata extracted from the encoded audio signal (102).

[0058] Reference dialogue loudness level (e.g. -31dB in "Line" mode) FS , -20dB in "RF" mode FS ) may be specified or established for a particular playback environment in the decoder (100). Additionally, alternatively, or optionally, in some embodiments, a user may be given control over setting or changing a reference dialogue loudness level in the decoder (100).

[0059] The DRC gain unit (114) can be configured to determine a dialogue loudness-related gain for the audio content to cause a change from an input dialogue loudness level to a reference dialogue loudness level as an output dialogue loudness level.

[0060] In some embodiments, the DRC gain unit (114) may be configured to handle peak levels in a particular playback environment in the decoder (100) and adjust the DRC gain to prevent clipping. In some embodiments, under the first approach, if the audio content extracted from the encoded audio signal (102) includes audio data elements for a reference multi-channel configuration with more channels than the particular speaker configuration in the decoder, a particular speaker configuration downmix may be performed from the reference multi-channel configuration before determining and processing the peak levels to prevent clipping. Additionally, optionally, or alternatively, in some embodiments, under the second approach, if the audio content extracted from the encoded audio signal (102) includes audio data elements for a reference multi-channel configuration with more channels than the particular speaker configuration in the decoder, a downmix formula (e.g., ITU stereo downmix, matrixed-surround compatible downmix, etc.) may be used to obtain peak levels for the particular speaker configuration in the decoder (100). The peak level may be adjusted to reflect the change from the input dialogue loudness level to a reference dialogue loudness level as the output dialogue loudness level. The maximum allowable gain without clipping (e.g., for an audio data block, for an audio data frame, etc.) may be determined based, at least in part, on the reciprocal of the peak level (e.g., multiplied by -1). In this way, an audio decoder based on the techniques described herein can be configured to accurately determine peak levels and apply clipping prevention specifically for decoder-side playback configurations. Neither the audio decoder nor the audio encoder need make assumptions about the worst-case scenario for a given decoder.In particular, the decoder in the first approach above can accurately determine peak levels and apply post-downmix clipping prevention without using the downmix equations, downmix channel gains, etc. (which are used under the second approach as described above).

[0061] In some implementations, the combination of adjustments to the dialogue loudness level and the DRC gain prevents peak level clipping even in the worst-case downmix (e.g., one that produces the largest peak level after the downmix, one that produces the largest downmix channel gain, etc.). However, in some other embodiments, the combination of adjustments to the dialogue loudness level and the DRC gain may not be sufficient to prevent peak level clipping. In these embodiments, the DRC gain may be replaced (e.g., capped) by the highest gain that prevents clipping at peak levels.

[0062] In some embodiments, the DRC gain unit (114) is configured to derive time constants (e.g., attack times, release times, etc.) from metadata extracted from the encoded audio signal (102). The DRC gains, time constants, maximum allowed gain, etc. may be used by the DRC gain unit (114) to perform DRC, gain smoothing, gain limiting, etc.

[0063] For example, the application of the DRC gain may be smoothed with a filter controlled by a time constant. The gain limiting operation may be implemented by a min() function that takes the smaller of the gain to be applied and the maximum allowed gain. Through this function, the gain (e.g., before limiting, DRC, etc.) may be replaced by the maximum allowed gain immediately, over a relatively short time interval, etc., thereby preventing clipping.

[0064] In some embodiments, the audio renderer (108) is configured to generate channel-specific (e.g., multi-channel) audio data (116) for the particular speaker configuration after applying a gain determined based on DRC, gain limiting, gain smoothing, etc. to the input audio data extracted from the encoded audio signal (102). The channel-specific audio data (118) may be used to drive speakers, headphones, etc. represented in the speaker configuration.

[0065] Additionally and / or optionally, in some embodiments, the decoder (100) may be configured to perform one or more other operations related to pre-processing, post-processing, rendering, etc., related to the input audio data.

[0066] The techniques described herein can be used with a variety of speaker configurations corresponding to a variety of different surround sound configurations (e.g., 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10-60 speaker configurations, 60+ speaker configurations, object signals or combinations of object signals, etc.) and a variety of different rendering environment configurations (e.g., movie theaters, parks, opera houses, concert halls, bars, homes, auditoriums, etc.).

[0067] 4. Audio Encoder 1B shows an exemplary encoder 150. The encoder 150 may include an audio content interface 152, a dialogue loudness analyzer 154, a DRC reference repository 156, an audio signal encoder 158, etc. The encoder 150 may be part of a broadcast system, an internet-based content server, an over-the-air network operator system, a film production system, etc.

[0068] In some embodiments, the audio content interface (152) is configured to receive audio content 160, audio content control inputs 162, etc., and to generate an encoded audio signal (e.g., 102) based at least in part on the audio content (160), some or all of the audio content control inputs (162), etc. For example, the audio content interface (152) may be used to receive audio content (160), audio content control inputs (162), etc. from a content creator, content provider, etc.

[0069] The audio content may be part of or in whole of an overall media data set that may include audio only, audiovisual, etc. The audio content (160) may include one or more of program portions, a program, several programs, one or more commercials, etc.

[0070] In some embodiments, the dialogue loudness analyzer (154) is configured to determine / establish one or more dialogue loudness levels for one or more portions of the audio content (152) (e.g., one or more programs, one or more commercials, etc.). In some embodiments, the audio content is represented by one or more collections of audio tracks. In some embodiments, dialogue audio content of the audio content is on a separate audio track. In some embodiments, at least a portion of the audio content is on an audio track containing non-dialogue audio content.

[0071] The audio content control inputs (162) may include some or all of the following: user control inputs, control inputs provided by systems / devices external to the encoder (150), control inputs from a content creator, control inputs from a content provider, etc. For example, a user, such as a mixing engineer, may provide / specify one or more dynamic range compression curve identifiers, which may be used to retrieve one or more dynamic range compression curves that best fit the audio content (160) from a data repository, such as the DRC reference repository (156).

[0072] In some embodiments, the DRC reference repository (156) is configured to store DRC reference parameter sets, etc. These DRC reference parameter sets may include definition data for one or more dynamic range compression curves, etc. In some embodiments, the encoder (150) may encode (e.g., concurrently, etc.) two or more dynamic range compression curves into the encoded audio signal (102). Zero, one, or more of the dynamic range compression curves may be standard-based, proprietary, customized, decoder-modifiable, etc. In one exemplary embodiment, both the dynamic range compression curves of FIGS. 2A and 2B may be encoded (e.g., concurrently, etc.) into the encoded audio signal (102).

[0073] In some embodiments, the audio signal encoder (158) can be configured to receive audio content from the audio content interface (152), a dialogue loudness level from the dialogue loudness analyzer (154), etc., retrieve one or more DRC reference parameter sets from the DRC reference repository (156), format the audio content into audio data blocks / frames, format the dialogue loudness level, the DRC reference parameter sets, etc. into metadata (e.g., metadata containers, metadata fields, metadata structures, etc.), encode the audio data blocks / frames and metadata into the encoded audio signal (102), etc.

[0074] Audio content to be encoded into an encoded audio signal as described herein may be received in one or more of a variety of source audio formats in one or more of a variety of ways, such as wirelessly, via a wired connection, through a file, via internet download, etc.

[0075] The encoded audio signals described herein can be part of an overall media data bitstream (e.g., for an audio broadcast, an audio program, an audiovisual program, an audiovisual broadcast, etc.). The media data bitstream can be accessed from a server, a computer, a media storage device, a media database, a media file, etc. The media data bitstream can be broadcast, transmitted, or received over one or more wireless or wired network links. The media data bitstream can be communicated through one or more intermediaries, such as a network connection, a USB connection, a wide area network, a local area network, a radio connection, an optical connection, a bus, a crossbar connection, a serial connection, etc.

[0076] Any of the components depicted (e.g., in Figures 1A, 1B, etc.) may be implemented as one or more processes and / or one or more integrated circuit circuits (e.g., ASICs, FPGAs, etc.) in hardware, software, or a combination of hardware and software.

[0077] 5. Dynamic Range Compression Curve 2A and 2B show exemplary dynamic range compression curves that can be used by the DRC gain unit (104) in the decoder (100) to derive DRC gains from input loudness levels. As shown, the dynamic range compression curves may be centered around a reference loudness level in the program to provide an appropriate overall gain for a particular playback environment. Exemplary definition data (e.g., in the metadata of the encoded audio signal 102) for the dynamic range compression curves (e.g., including, but not limited to, boost ratios, cut ratios, attack times, release times, etc.) are shown in the table below. Here, each profile in the multiple profiles (e.g., film standard, film light, music standard, music light, speech, etc.) represents a particular playback environment (e.g., in the decoder 100).

[0078] [Table 1] Some embodiments may use dB SPL or dB FS Loudness level expressed in dB and SPL The DRC gain may be expressed in dB. SPLA different loudness representation (e.g., sones) that has a nonlinear relationship with loudness level may be implemented, and the compression curve used in the DRC gain calculation may then be transformed to be described using the different loudness representation (e.g., sones).

[0079] 6. DRC Gain, Gain Limiting and Gain Smoothing 3 shows exemplary processing logic for determining / calculating a combined DRC and limiting gain. The processing logic may be implemented by a decoder (100), an encoder (150), etc. For illustrative purposes only, a DRC gain unit (e.g., 114) in a decoder (e.g., 100) may be used to implement the processing logic.

[0080] A DRC gain for a portion of audio content (e.g., one or more audio data blocks, audio data frames, etc.) may be calculated using a loudness level determined from the portion of audio content. The loudness level may first be adjusted with respect to (e.g., relative to) a dialogue loudness level (e.g., of a program of which the audio content is a part) in metadata extracted from the encoded audio signal (102). In the example shown in Figure 3, the difference between the loudness level of the portion of audio content and the dialogue loudness level ("dialnorm") may be used as input to find a DRC gain from a selected dynamic range compression curve.

[0081] To prevent clipping of the output audio data elements in that particular playback environment, the DRC gain unit (114) may be configured to handle peak levels in a particular playback scenario (e.g., specific to a particular combination of encoded audio signal 102 and playback environment in decoder 100), which may be one of a variety of possible playback scenarios (e.g., a multi-channel scenario, a downmix scenario, etc.).

[0082] In some embodiments, individual peak levels for individual portions of the audio content (e.g., audio data blocks, several audio data blocks, audio data frames, etc.) at a particular time resolution may be provided as part of the metadata extracted from the encoded audio signal (102).

[0083] In some embodiments, the DRC gain unit (114) can be configured to determine peak levels in these scenarios and adjust the DRC gain if necessary. During the DRC gain calculation, a parallel process may be used by the DRC gain unit (114) to determine peak levels of the audio content. For example, the audio content may be encoded for a reference multi-channel configuration with more channels than the channels of a particular speaker configuration used by the decoder (100). The audio content for the more channels of the reference multi-channel configuration may be converted to downmixed audio data (e.g., an ITU stereo downmix, a matrixed-surround compatible downmix, etc.) to derive fewer channels for the particular speaker configuration in the decoder (100). In some embodiments, under the first approach, downmixing from the reference multi-channel configuration to the particular speaker configuration may be performed before determining and processing peak levels to prevent clipping. Additionally, optionally, or alternatively, in some embodiments, under the second approach, downmix channel gains associated with downmixing audio content may be used as part of the input for adjusting, deriving, calculating, etc., peak levels for that particular speaker configuration. In an example embodiment, the downmix channel gains may be derived based at least in part on one or more downmix equations used to perform a downmix operation from a reference multi-channel configuration to a particular speaker configuration in a playback environment in the decoder (100).

[0084] In some media applications, the reference dialogue loudness level (e.g., -31 dB in "Line" mode) FS , -20dB in "RF" mode FS) may be specified or assumed for a particular playback environment in the decoder (100). In some embodiments, the user may be given control over setting or changing the reference dialogue loudness level in the decoder (100).

[0085] A dialogue loudness related gain may be applied to the audio content to adjust the (e.g. output) dialogue loudness level to a reference dialogue loudness level. To reflect this adjustment, the peak level should be adjusted accordingly. In one example, the (input) dialogue loudness level is -23dB. FS Reference dialogue loudness level is -31dB FS In "Line" mode, the adjustment to the (input) dialogue loudness level is -8dB to produce an output dialogue loudness level of the reference dialogue loudness level. In this "Line" mode, the adjustment to the peak level is also -8dB, the same as the adjustment to the dialogue loudness level. FS In "RF" mode, the adjustment to the (input) dialogue loudness level is 3 dB to result in an output dialogue loudness level of the reference dialogue loudness level. In this "RF" mode, the adjustment to the peak level is also 3 dB, the same as the adjustment to the dialogue loudness level.

[0086] The peak level and the sum of the difference between the reference dialogue loudness level (denoted "dialref") and the dialogue loudness level in the metadata from the encoded audio signal (102) ("dialnorm") may be used as input to calculate the maximum (e.g., allowed) gain for the DRC gain. The adjusted peak level (0 dB FS (relative to clipping level) dB FS, so the maximum allowable gain without clipping (e.g., for the current audio data block, for the current audio data frame, etc.) is simply the inverse of the adjusted peak level (e.g., multiplied by -1).

[0087] In some embodiments, the peak level may be higher than the clipping level (0 dB), even though the dynamic range compression curve from which the DRC gain is derived is designed to cut some loud sounds. FS In some embodiments, the combination of the dialogue loudness level and adjustments to the DRC gain prevents clipping at peak levels, even in the worst-case downmix (e.g., one that produces the largest downmix channel gain). However, in other embodiments, the combination of the dialogue loudness level and adjustments to the DRC gain may not be sufficient to prevent clipping at peak levels. In these embodiments, the DRC gain may be replaced (e.g., capped) by the highest gain that prevents clipping at peak levels.

[0088] In some embodiments, the DRC gain unit (114) is configured to derive time constants (e.g., attack time, release time, etc.) from metadata extracted from the encoded audio signal (102). These time constants may or may not vary with one or more of the dialogue loudness level or the current loudness level of the audio content. The DRC gain, time constant, and maximum gain retrieved from the dynamic range compression curve may be used to perform gain smoothing and limiting operations.

[0089] In some embodiments, the DRC gain, which may be potentially gain-limited, does not exceed the maximum peak loudness level in a particular playback environment. The static DRC gain derived from the loudness level may be smoothed with a filter controlled by a time constant. The limiting operation may be implemented by one or more min() functions, through which the DRC gain (e.g., before limiting) may be immediately replaced by the maximum allowed gain, such as over a relatively short time interval, thereby preventing clipping. The DRC algorithm may be configured to smoothly release from the clipping gain to a lower gain as the peak level of the incoming audio content transitions from above the clipping level to below the clipping level.

[0090] One or more different (e.g., real-time, two-pass, etc.) implementations may be used to perform the DRC gain determination / calculation / application shown in FIG. 3 . For illustrative purposes only, adjustments to the dialogue loudness level, the (e.g., static, etc.) DRC gain, time-dependent gain variations due to smoothing, gain clipping due to limiting, etc. have been described as a combined gain from the DRC algorithm above. However, in various embodiments, other approaches may be used to apply gain to audio content for dialogue loudness level control (e.g., between different programs), dynamic range control (e.g., for different portions of the same program), clipping prevention, gain smoothing, etc. For example, some or all of the adjustments to the dialogue loudness level, the (e.g., static, etc.) DRC gain, time-dependent gain variations due to smoothing, gain clipping due to limiting, etc. may be applied partially / individually, applied serially, applied in parallel, applied partially in serial and partially in parallel, etc.

[0091] 7. Input Smoothing and Gain Smoothing In addition to DRC gain smoothing, various embodiments may implement other smoothing processes under the techniques described herein. In one example, input smoothing may be used, where input audio data extracted from the encoded audio signal (102) may be smoothed using, for example, a simple single-pole smoothing filter, to obtain a spectrum of specific loudness levels with better temporal characteristics (e.g., smoother in time, less spikes in time, etc.) than the spectrum of specific loudness levels without input smoothing.

[0092] In some embodiments, different smoothing processes described herein may use different time constants (e.g., 1 second, 4 seconds, etc.). In some embodiments, two or more smoothing processes may use the same time constant. In some embodiments, the time constants used in the smoothing processes described herein may be frequency dependent. In some embodiments, the time constants used in the smoothing processes described herein may be frequency independent.

[0093] One or more smoothing processes may be connected to a reset process that supports automatic or manual resetting of the one or more smoothing processes. In some embodiments, when a reset occurs in the reset process, the smoothing process may speed up the smoothing operation by switching or transitioning to a smaller time constant. In some embodiments, when a reset occurs in the reset process, the memory of the smoothing process may be reset to a value. This value may be the last input sample to the smoothing process.

[0094] 8. DRC across multiple frequency bands In some embodiments, specific loudness levels in specific frequency bands can be used to derive corresponding DRC gains in those specific frequency bands. However, this can lead to changes in timbre, because those specific loudness levels can vary significantly in different bands and therefore suffer from different DRC gains, even when the broadband loudness level across all frequency bands remains constant.

[0095] In some embodiments, rather than applying DRC gains that vary with individual frequency bands, DRC gains that do not vary with frequency bands but instead vary with time are applied. The same time-varying DRC gain is applied across all frequency bands. The time-averaged DRC gain of the time-varying DRC gain may be set equal to a static DRC gain derived from the selected dynamic range compression curve based on broadband, wideband, and / or overall loudness levels across a broadband (or wideband) range or multiple frequency bands. As a result, changes in tonal effects that may be caused by applying different DRC gains to different frequency bands in other approaches can be avoided.

[0096] In some embodiments, the DRC gain in each frequency band is controlled using a broadband (or wideband) DRC gain determined based on the broadband (or wideband) loudness level. The DRC gain in each frequency band may operate around the broadband (or wideband) DRC found in the dynamic range compression curve based on the broadband (or wideband) loudness level. Thus, the DRC gain in each frequency band time-averaged over a time interval (e.g., longer than 5.3 ms, 20 ms, 50 ms, 80 ms, 100 ms, etc.) is the same as the broadband (or wideband) level shown in the dynamic range compression curve. In some embodiments, loudness level fluctuations over short time intervals relative to the time interval that deviate from the time-averaged DRC gain are tolerable between channels and / or frequency bands. This approach ensures application of the correct multi-channel and / or multi-band time-averaged DRC gains shown in the dynamic range compression curve and prevents the DRC gains in short time intervals from deviating too greatly from such time-averaged DRC gains shown in the dynamic range compression curve.

[0097] 9. Volume adjustment in the loudness range Applying linear processing for volume adjustment to an audio excitation signal under other approaches that do not implement the techniques described herein can make low audible signal levels inaudible (e.g., below the frequency-dependent hearing threshold of the human auditory system).

[0098] Under the techniques described in this paper, volume adjustment of audio content is performed in the physical domain (e.g., dB SPLIn some embodiments, the loudness levels of all bands are scaled by the same factor in the loudness domain to maintain the perceptual quality and / or integrity of the loudness level relationships among all bands at all volume levels. The volume adjustments described herein based on setting and adjusting gains in the loudness domain may be translated back to and implemented through nonlinear processing in the physical domain (or in a digital domain representing the physical domain) that applies different scaling factors to the audio excitation signal in different frequency bands. The nonlinear processing in the physical domain, translated from the volume adjustments in the loudness domain under the techniques described herein, attenuates or boosts the loudness level of the audio content with a DRC gain that prevents most or all of the lower audible levels in the audio content from being inaudible. In some embodiments, the difference in loudness levels between loud and soft sounds in a program is reduced—but not perceptually eliminated—using these DRC gains to keep low audible signal levels above the hearing threshold of the human auditory system. In some embodiments, to maintain similarities in spectral perception and perceived timbre over a large range of volume levels, frequencies or frequency bands with excitation signal levels near the hearing threshold are attenuated less at low volume levels and are therefore perceptually audible.

[0099] The techniques described herein may implement transformations (e.g., back and forth transformations) between signal levels, gains, etc. in the physical domain (or a digital domain representing the physical domain) and loudness levels, gains, etc. in the loudness domain. These transformations may be based on forward and inverse versions of one or more nonlinear functions (e.g., mappings, curves, piecewise linear segments, lookup tables, etc.) constructed based on models of the human auditory system.

[0100] 10. Gain Profile with Differential Gain In some embodiments, an audio encoder described herein (such as 150) is configured to provide profile-related metadata to a downstream audio decoder, for example, the profile-related metadata may be carried in the encoded audio signal as part of the audio-related metadata along with the audio content.

[0101] The profile-related metadata described herein includes, but is not limited to, definition data for a plurality of gain profiles. One or more first gain profiles (referred to as one or more default gain profiles) in the plurality of gain profiles are represented by one or more corresponding DRC curves (referred to as one or more default DRC curves). The definition data is included in the profile-related metadata. One or more second gain profiles (referred to as one or more non-default gain profiles) in the plurality of gain profiles are represented by one or more corresponding sets of differential gains with respect to the one or more default DRC curves. The definition data is included in the profile-related metadata. More specifically, a default DRC curve (e.g., in the profile-related metadata) can be used to represent a default gain profile, and a set of differential gains (e.g., in the profile-related metadata) with respect to a default gain profile can be used to represent a non-default gain profile.

[0102] In some embodiments, a set of differential gains representing a non-default gain profile relative to a default DRC curve representing a default gain profile includes a gain difference (or gain adjustment) between the set of non-differential (e.g., non-default, etc.) gains generated for the non-default gain profile and the set of non-differential (e.g., default, etc.) gains generated for the default gain profile. Examples of non-differential gains include, but are not limited to, null gain, DRC gain or attenuation, gain or attenuation for dialogue normalization, gain limiting, gain smoothing, etc. The gains (e.g., non-differential gains, differential gains, etc.) described herein may be time-dependent and may have values ​​that change over time.

[0103] To generate a set of non-differential gains for a gain profile (e.g., a default gain profile, a non-default gain profile, etc.), an audio encoder described herein may perform a set of gain-profile-specific gain-generation operations, which may include DRC operations, gain-limiting operations, gain-smoothing operations, etc. This includes, but is not limited to, any of the following: (1) globally applicable to all gain profiles; (2) specific to one or more, but not all, gain profiles, specific to one or more default DRC curves; (3) specific to one or more non-default DRC curves; (4) specific to a corresponding (e.g., default, non-default, etc.) gain profile; (5) relating to one or more algorithms, curves, functions, operations, parameters, etc. that exceed the limits of parameterization supported by a media coding format, media standard, media proprietary specification, etc.; or (6) relating to one or more algorithms, curves, functions, operations, parameters, etc. that are not yet commonly implemented in commercially available audio decoding devices.

[0104] In some embodiments, the audio decoder (150) can be configured to determine a set of differential gains for the audio content (152) based, at least in part, on a default gain profile represented by a default DRC curve (e.g., by definition data in profile-related metadata of the encoded audio signal) and a non-default gain profile that differs from the default gain profile, and include the set of differential gains as part of the profile-related metadata in the encoded audio signal as a representation of the non-default gain profile (e.g., relative to the default DRC curve). The set of differential gains extracted from the profile-related metadata in the encoded audio signal relative to the default DRC curve can be used by a receiving audio decoder to efficiently and consistently perform a gain operation (or attenuation operation) in a playback environment or scenario for a particular gain profile represented by the set of differential gains relative to the default DRC curve. This allows the receiving audio decoder to apply gain or attenuation for that particular gain profile without requiring the receiving audio decoder to implement a set of gain-generating operations. To generate the gain or attenuation, a set of gain generation operations can be implemented in the audio encoder (150).

[0105] In some embodiments, one or more sets of differential gains may be included in the profile-related metadata by the audio encoder (150). Each of the one or more sets of differential gains may be derived from a corresponding non-default gain profile in one or more non-default gain profiles relative to a corresponding default gain profile in one of the one or more default gain profiles. For example, a first set of differential gains in the one or more sets of differential gains may be derived from a first non-default gain profile relative to a first default gain profile, while a second set of differential gains in the sets of differential gains may be derived from a second non-default gain profile relative to a second default gain profile.

[0106] In some embodiments, the first set of differential gains comprises a first gain difference (or gain adjustment) determined between a first set of non-differential non-default gains generated based on the first non-default gain profile and a first set of non-differential default gains generated based on the first default gain profile, while the second set of differential gains comprises a second gain difference determined between a second set of non-differential non-default gains generated based on the second non-default gain profile and a second set of non-differential default gains generated based on the second default gain profile.

[0107] The first default gain profile and the second default gain profile may be the same (e.g., represented by the same default DRC curve with the same set of gain-generating operations), or may be different (e.g., represented by different default DRC curves, represented by a default DRC with a different set of gain-generating operations, etc.). In various embodiments, additionally, optionally, or alternatively, the first non-default gain profile may or may not be the same as the second non-default gain profile.

[0108] The profile-related metadata generated by the audio encoder (150) may carry one or more specific flags, indicators, data fields, etc. to indicate the presence of one or more sets of differential gains for one or more corresponding non-default gain profiles. The profile-related data may also include preference flags, indicators, data fields, etc. to indicate which non-default gain profiles are preferred for rendering the audio content in a particular playback environment or scenario.

[0109] In some embodiments, an audio decoder (e.g., 100) described herein is configured to decode (e.g., multi-channel) audio content from an encoded audio signal (102), such as extracting a dialogue loudness level (e.g., "dialnorm") from loudness metadata delivered with the audio content.

[0110] In some embodiments, the audio decoder (e.g., 100) is configured to perform at least one set of gain generation operations for a gain profile, such as the first default profile, the second default profile, etc. For example, the audio decoder (100) can decode an encoded audio signal (102) having a dialogue loudness level (e.g., "dialnorm"); perform a set of gain generation operations to obtain a set of non-differential default gains (or attenuations) for a default gain profile represented by a default DRC curve whose definition data can be extracted by the audio decoder (100) from the encoded audio signal (102); apply the set of non-differential default gains (e.g., the difference between a reference loudness level and "dialnorm") for the default gain profile during decoding to align / adjust an output dialogue loudness level of a sound output to the reference loudness level; etc.

[0111] Additionally, optionally, or alternatively, in some embodiments, the audio decoder (100) is configured to extract at least one set of differential gains from the encoded audio signal (102), the set of differential gains representing a non-default gain profile relative to a default DRC curve as discussed above as part of the metadata delivered with the audio content. In some embodiments, the profile-related metadata includes one or more different sets of differential gains, each of the one or more different sets of differential gains representing a non-default gain profile relative to a respective default DRC curve representing a default gain profile. The presence of a DRC curve or a set of differential gains in the profile-related metadata may be indicated by one or more flags, indicators, or data fields carried in the profile-related metadata.

[0112] In response to determining that the one or more sets of differential gains exist, the audio decoder (100) can determine / select, from the one or more different sets of differential gains, a set of differential gains that corresponds to a particular non-default gain profile. The audio decoder (100) can further be configured to identify—e.g., among definition data for one or more different default DRC curves in the profile-related metadata—a default DRC curve based on which the set of differential gains represents the particular gain profile.

[0113] In some embodiments, the audio decoder (100) is configured to perform a set of gain generation operations to obtain a set of non-differential default gains (or attenuation) for the default gain profile. The set of gain generation operations performed by the audio decoder (100) to obtain the set of non-differential default gains based on a default DRC curve may include one or more operations related to one or more of a standard, a proprietary specification, etc. In some embodiments, the audio decoder (100) is configured to: generate a set of non-differential non-default gains for the particular non-default gain profile based on the set of differential gains whose definition data is extracted from profile-related metadata and the set of non-differential default gains generated by the set of gain generation operations based on a default DRC curve; apply the set of non-differential non-default gains (e.g., the difference between a reference loudness level and a "dialnorm") for the default gain profile during decoding to align / adjust an output dialogue loudness level of a sound output to a reference loudness level;

[0114] In some embodiments, the audio decoder (100) can perform gain-related operations for one or more gain profiles. The audio decoder (100) can be configured to determine and perform gain-related operations for a particular gain profile based on one or more factors. These factors may include, but are not limited to, one or more of: user input specifying a preference for a particular user-selected gain profile; user input specifying a preference for a system-selected gain profile; capabilities of the particular speaker or audio channel configuration used by the audio decoder (100); capabilities of the audio decoder (100); availability of profile-related metadata for the particular gain profile; any encoder-generated preference flags for the gain profile; etc. In some embodiments, the audio decoder (100) may implement one or more procedural rules, seek further user input, etc., to determine or select a particular gain profile when there is a conflict between these factors.

[0115] 11. Additional Actions Related to Gains Under the techniques described herein, other processing such as dynamic equalization, noise compensation, etc. can also be performed in the loudness (e.g., perceptual) domain rather than the physical domain (or a digital domain representing the physical domain).

[0116] In some embodiments, the gains from some or all of the various processes, such as DRC, equalization noise compensation, clip prevention, gain smoothing, etc., may be combined into the same gain in the loudness domain and / or may be applied in parallel. In other embodiments, the gains from some or all of the various processes, such as DRC, equalization noise compensation, clip prevention, gain smoothing, etc., may be separate gains in the loudness domain and / or may be applied at least partially in series. In other embodiments, the gains from some or all of the various processes, such as DRC, equalization noise compensation, clip prevention, gain smoothing, etc., may be applied sequentially.

[0117] 12. Specific and Broadband Loudness Levels One or more audio processing elements, units, components, etc., such as transmit filters, auditory filter banks, synthesis filter banks, short-time Fourier transforms, etc., may be used by an encoder or decoder to perform the audio processing operations described herein.

[0118] In some embodiments, one or more transmission filters that model the filtering of the outer and middle ear of the human auditory system may be used to filter the incoming audio signal (e.g., the encoded audio signal 102, audio content from a content provider, etc.). In some embodiments, an auditory filter bank may be used to model the frequency selectivity and frequency spread of the human auditory system. The excitation signal level from some or all of these filters may be determined / calculated and smoothed with a frequency-dependent time constant that is shorter at higher frequencies to model the integration of energy in the human auditory system. A nonlinear function (e.g., relationship, curve, etc.) between the excitation signal and the specific loudness level may then be used to obtain a frequency-dependent specific loudness level profile. A broadband (or wideband) loudness level can be obtained by integrating the specific loudness over frequency bands.

[0119] A straightforward summation / integration of the specific loudness level (e.g., using equal weights for all frequency bands) may work well for broadband signals. However, such an approach may underestimate the (e.g., perceptual) loudness level for narrowband signals. In some embodiments, the specific loudness levels at different frequencies or in different frequency bands are given different weights.

[0120] In some embodiments, the auditory filterbank and / or transmission filter as described above may be replaced by one or more short-time Fourier transforms (STFTs). The responses of the transmission filters and auditory filterbank may be applied in the fast Fourier transform (FFT) domain. In some embodiments, one or more inverse transmission filters are used, for example, when one or more (e.g., forward) transmission filters are used in or before the transformation from the physical domain (or a digital domain representing the physical domain) to the loudness domain. In some embodiments, no inverse transmission filters are used, for example, when an STFT is used instead of the auditory filterbank and / or transmission filters. In some embodiments, the auditory filterbank is omitted; instead, one or more quadrature mirror filters (QMFs) are used. In these embodiments, the diffusion effect of the basilar membrane in models of the human auditory system may be omitted without significantly affecting the audio processing operations described herein.

[0121] Under the techniques described herein, different embodiments may use different numbers of frequency bands (e.g., 20 frequency bands, 40 frequency bands, etc.) Additionally, optionally, or alternatively, different bandwidths may be used in different embodiments.

[0122] 13. Individual Gains for Individual Subsets of Channels In some embodiments, when a particular speaker configuration is a multi-channel configuration, the overall loudness level may be obtained by first summing the excitation signals of all channels before conversion from the physical domain (or a digital domain representing the physical domain) to the loudness domain. However, applying the same gain to all channels in a particular speaker configuration may not preserve the spatial balance (balance in terms of relative loudness levels between different channels, etc.) between the different channels of that particular speaker configuration.

[0123] In some embodiments, to preserve spatial balance so that relative perceptual loudness levels between different channels can be optimally or correctly maintained, respective loudness levels and corresponding gains obtained based on the respective loudness levels may be determined or calculated for each channel. In some embodiments, the corresponding gains obtained based on the respective loudness levels are not equal to the same overall gain. For example, some or all of the corresponding gains may each be equal to the overall gain plus a small (e.g., channel-specific) correction.

[0124] In some embodiments, to preserve spatial balance, the respective loudness levels and corresponding gains obtained based on the respective loudness levels may be determined or calculated for each subset of channels. In some embodiments, the corresponding gains obtained based on the respective loudness levels are not equal to the same overall gain. For example, some or all of the corresponding gains may each be equal to the overall gain plus a small (e.g., channel-specific) correction. In some embodiments, a subset of channels may include two or more channels that are a true subset of all channels for that particular speaker configuration (e.g., a subset of channels including left front, right front, and low-frequency effects (LFE), a subset of channels including left surround and right surround, etc.). The audio content for a subset of channels may form a submix of the overall mix carried in the encoded audio signal (102). The channels within a submix may have the same gain applied to them.

[0125] In some embodiments, to generate an actual loudness (e.g., as actually perceived) from a particular speaker configuration, the signal level in the digital domain is converted to a corresponding physical (e.g., dB) in the physical domain represented by the digital domain. SPL One or more calibration parameters may be used to relate the spatial pressure (e.g., spatial pressure) levels to the sound pressure levels, which may be given values ​​specific to the physical sound installation in a particular speaker configuration.

[0126] 14. Auditory Scene Analysis In some embodiments, the encoders described herein may implement computer-based auditory scene analysis (ASA) to detect auditory event boundaries in audio content (e.g., as encoded in encoded audio signal 102), generate one or more ASA parameters, and format the one or more ASA parameters as part of the encoded audio signal (e.g., 102) that is delivered to a downstream device (e.g., decoder 100). The ASA parameters may include, but are not limited to, the location of auditory event boundaries, values ​​of auditory event certainty indicators (described further below), etc.

[0127] In some implementations, the location (e.g., in time) of auditory event boundaries may be indicated in metadata encoded within the encoded audio signal (102). Additionally, optionally, or alternatively, the location (e.g., in time) of auditory event boundaries may be indicated (e.g., using a flag, data field, etc.) in the audio data block and / or frame in which the location of the auditory event boundary is detected.

[0128] As used herein, an auditory event boundary refers to the point where a preceding auditory event ends and / or a subsequent auditory event begins. Each auditory event occurs between two successive auditory event boundaries.

[0129] In some embodiments, the encoder (150) is configured to detect auditory event boundaries by a difference in a specific loudness spectrum between two consecutive (e.g., temporally) audio data frames, where each specific loudness spectrum may comprise an unsmoothed loudness spectrum calculated from a corresponding one of the consecutive audio data frames.

[0130] In some embodiments, the specific loudness spectrum N[b,t] is calculated as the normalized specific loudness spectrum N[b,t], as shown in the following equation: NORM It may be normalized to obtain [b,t].

[0131] N NORM [b,t]=N[b,t] / max b {N[b,t]} (1) where b denotes the bandwidth, t denotes the time or audio frame index, and max b {N[b,t]} is the maximum specific loudness level across all frequency bands.

[0132] The normalized specific loudness spectra are subtracted from each other and used to derive the sum of absolute differences D[t], as shown in the following equation:

[0133] D[t]=Σ b |N NORM [b,t]-N NORM [b,t-1]| (2) The sum of absolute differences is mapped to an auditory event certainty index with a value range from 0 to 1 as follows:

[0134]

number

[0135] In some embodiments, the encoder (150) determines whether D[t] (e.g., at a particular t) is D min and configured to detect an auditory event boundary (e.g., at said particular t) when t exceeds

[0136] In some embodiments, the decoders described herein (e.g., 100) extract ASA parameters from an encoded audio signal (e.g., 102) and use the ASA parameters to prevent unintentional boosting of soft sounds and / or unintentional cutting of loud sounds, which would cause perceptual distortion of auditory events.

[0137] The decoder (100) may be configured to mitigate or prevent unintended distortion of auditory events by ensuring that gain is more constant within the auditory event and constraining many of the gain changes to the vicinity of auditory event boundaries. For example, the decoder (100) may be configured to use a relatively small time constant (e.g., comparable to or shorter than the minimum duration of auditory events) in response to gain changes in the attack (e.g., loudness level increase) of an auditory event boundary. Thus, gain changes in the attack can be implemented relatively quickly by the decoder (100). On the other hand, the decoder (100) may be configured to use a relatively long time constant compared to the duration of the auditory event in response to gain changes in the release (e.g., loudness level decrease) of an auditory event. Thus, gain changes in the release can be implemented relatively slowly by the decoder (100), so that sounds that should be perceived as constant or that decay gradually may not be auditorily or perceptually disruptive. A fast response in the attack of auditory event boundaries and a slow response in the release of auditory events allows for a fast perception of the arrival of auditory events while preserving the perceptual quality and / or integrity between auditory events, such as piano chords, which contain loud and soft sounds linked by specific loudness-level relationships and / or specific time relationships.

[0138] In some embodiments, auditory event boundaries indicated by auditory events and ASA parameters are used by the decoder (100) to control gain changes in one, two, some, or all of the channels in a particular speaker configuration in the decoder (100).

[0139] 15. Loudness Level Transition A loudness level transition may occur, for example, between two programs, between a program and a loud commercial, etc. In some embodiments, the decoder (100) is configured to maintain a histogram of instantaneous loudness levels based on past audio content (e.g., received from the encoded audio signal 102 over the past four seconds). Over the time interval from before the loudness level transition to after the loudness level transition, two regions of increased probability may be recorded in the histogram. One of the regions is centered around the previous loudness level, while the other of the regions is centered around the new loudness level.

[0140] The decoder (100) may dynamically determine a smoothed loudness level as audio content is processed and determine a corresponding bin of a histogram (e.g., an instantaneous loudness level bin containing the same value as the smoothed loudness level) based on the smoothed loudness level. The decoder (100) may be further configured to compare the probability in the corresponding bin with a threshold (e.g., 6%, 7%, 7.5%, etc.), where the total area of ​​the histogram curve (e.g., the sum of all bins) represents a 100% probability. The decoder may be configured to detect the occurrence of a loudness level transition by determining the probability that the corresponding bin falls below the threshold. In response, the decoder (100) may be configured to select a relatively small time constant to adapt relatively quickly to the new loudness level. As a result, the duration of a loud (or soft) onset within a loudness level transition can be shortened.

[0141] In some embodiments, the decoder (100) uses a silence / noise gate to prevent low instantaneous loudness levels from entering the histogram and resulting in high probability bins in the histogram. Additionally, optionally, or alternatively, the decoder (100) may be configured to use the ASA parameter to detect auditory events to be included in the histogram. In some embodiments, the decoder (100) uses a time-dependent value of the time-averaged auditory event certainty index (ASA) to detect auditory events to be included in the histogram.

number

number

number

[0142] In some embodiments, for (e.g., instantaneous) loudness levels that are allowed to be included in the histogram (e.g., have corresponding A[t] values ​​above the histogram inclusion threshold), the loudness levels are assigned weights that are equal to or proportional to the time-dependent values ​​of the time-averaged auditory event certainty index A[t] concurrently with those loudness levels. As a result, loudness levels that are close to auditory event boundaries have more influence on the histogram (e.g., have relatively larger A[t] values) than other loudness levels that are not close to auditory event boundaries.

[0143] 16. Reset In some embodiments, an encoder (e.g., 150) described herein is configured to detect a reset event and include an indication of the reset event in the encoded audio signal (e.g., 102). In a first example, the encoder (150) detects a reset event in response to determining that a continuous period of relative silence (e.g., 250 milliseconds, configurable by the system and / or user) occurs. In a second example, the encoder (150) detects a reset event in response to determining that a large momentary drop in excitation level occurs across all frequency bands. In a third example, the encoder is given input (e.g., user input, system-controlled metadata, etc.) where a content transition (e.g., program start / end, scene change, etc.) occurs that requires a reset.

[0144] In some embodiments, the decoders described herein (such as 100) implement a reset mechanism that can be used to speed up instantaneous gain smoothing, which may be useful and invoked when switching between channels and audiovisual inputs occurs.

[0145] In some embodiments, the decoder (100) can be configured to determine whether a reset event occurs by determining whether a continuous period of relative silence (e.g., 250 milliseconds configurable by the system and / or user) occurs, whether a large momentary drop in excitation level across all frequency bands occurs, etc.

[0146] In some embodiments, the decoder (100) is configured to determine that a reset event has occurred in response to receiving an indication (e.g., a reset event) provided in the encoded audio signal (102) by an upstream encoder (e.g., 150).

[0147] The reset mechanism may be configured to issue a reset when the decoder (100) determines that a reset event has occurred. In some embodiments, the reset mechanism is configured to use a more aggressive cut behavior of the DRC compression curve to prevent hard starts (e.g., for loud programs / channels / audiovisual sources, etc.). Additionally, optionally, or alternatively, the decoder (100) may be configured to implement safeguards to recover gracefully when the decoder (100) detects that a reset has been falsely triggered.

[0148] 17. Gain Provided by the Encoder In some embodiments, the audio decoder may be configured to calculate one or more sets of gains (e.g., DRC gains, etc.) for individual portions of audio content to be encoded into the encoded audio signal (e.g., audio data blocks, audio data frames, etc.). The sets of gains generated by the audio encoder may include: a first set of gains including a single broadband (or wideband) gain for all channels (e.g., left front, right front, low frequency effects or LFE, center, left surround, right surround, etc.); a second set of gains including individual broadband (or wideband) gains for individual subsets of channels; a third set of gains including individual broadband (or wideband) gains for individual subsets of channels and for each of a first number (e.g., two) individual bands (e.g., two bands in each channel); a fourth set of gains including individual broadband (or wideband) gains for individual subsets of channels and for each of a second number (e.g., four) individual bands (e.g., four bands in each channel); The subset of channels described herein may be one of a subset including left front, right front and LFE channels, a subset including a center channel, a subset including left surround and right surround channels, etc.

[0149] In some embodiments, an audio encoder is configured to transmit one or more portions of audio content (e.g., audio data blocks, audio data frames, etc.) and one or more sets of gains calculated for the one or more portions of audio content in a time-synchronous manner. An audio decoder receiving the one or more portions of audio content can select and apply a set of gains from the one or more sets of gains with little or no delay. In some embodiments, the audio encoder can implement a subframing technique in which the one or more sets of gains are carried in one or more subframes (e.g., using differential encoding, etc.) as shown in FIG. 4. In one example, subframes may be encoded within the audio data block or audio data frame for which the gains are calculated. In another example, subframes may be encoded within audio data blocks or audio data frames that precede the audio data block or audio data frame for which the gains are calculated. In another non-limiting example, subframes may be encoded into audio data blocks or frames within a certain time from the audio data block or frame for which the gains are calculated. In some embodiments, Huffman and differential coding may be used to populate and / or compress the subframes that carry those sets of gains.

[0150] 18. Exemplary System and Process Flows FIG. 5 illustrates an exemplary codec system in a non-limiting, example embodiment. A content creator, which may be a processing unit within an audio encoder such as 150, is configured to provide audio content (“audio”) to an encoder unit (“NGC Encoder”). The encoder unit formats the audio content into audio data blocks and / or frames and encodes the audio data blocks and / or frames into an encoded audio signal. The content creator is also configured to establish / generate one or more dialogue loudness levels (“dialnorms”) and one or more dynamic range compression curve identifiers (“compression curve IDs”) for one or more programs, commercials, etc. in the audio content. The content creator may determine the dialogue loudness levels from one or more dialogue audio tracks in the audio content. The dynamic range compression curve identifiers may be selected based, at least in part, on user input, system configuration parameters, etc. A content creator may be a human (e.g., artist, audio engineer, etc.) who uses tools to generate audio content and dialnorms.

[0151] Based on the dynamic range compression curve identifier, the encoder (150) generates one or more DRC parameter sets, including, but not limited to, corresponding reference dialogue loudness levels ("reference levels") for multiple playback environments supported by the one or more dynamic range compression curves. These DRC parameter sets may be encoded in the metadata of the encoded audio signal, such as in-band with the audio content, out-of-band with the audio content, etc. Operations such as compression, format multiplexing ("MUX"), etc. may be performed as part of generating an encoded audio signal that can be delivered to an audio decoder, such as 100. The encoded audio signal may be encoded with syntax that supports carrying audio data elements, DRC parameter sets, reference loudness levels, dynamic range compression curves, functions, lookup tables, Huffman codes used in compression, subframes, etc. In some embodiments, the syntax allows an upstream device (e.g., encoder, decoder, transcoder, etc.) to transmit gains to a downstream device (e.g., decoder, transcoder, etc.). In some embodiments, the syntax used to encode data into and / or decode data from an encoded audio signal is configured to support backward compatibility, so that devices that rely on gains calculated by upstream devices may optionally continue to do so.

[0152] In some embodiments, the encoder (150) calculates two or more sets of gains (e.g., gain smoothing using an appropriate reference dialogue loudness level, DRC gains, etc.) for the audio content. These sets of gains may be provided in metadata encoded into an audio signal encoded along with the audio content, along with the one or more dynamic range compression curves. A first set of gains may correspond to broadband (or wideband) gains for all channels in a speaker configuration or profile (e.g., default). A second set of gains may correspond to broadband (or wideband) gains for each of all channels in a speaker configuration or profile. A third set of gains may correspond to broadband (or wideband) gains for each of two bands in each of all channels in a speaker configuration or profile. A fourth set of gains may correspond to broadband (or wideband) gains for each of four bands in each of all channels in a speaker configuration or profile. In some embodiments, the set of gains calculated for a speaker configuration may be transmitted in metadata along with the (e.g., parameterized) dynamic range compression curve for the speaker configuration. In some embodiments, a set of gains calculated for a speaker configuration may replace the (e.g., parameterized) dynamic range compression curve for that speaker configuration in the metadata. Additional speaker configurations or profiles may be supported under the techniques described herein.

[0153] The decoder (100) is configured to extract audio data blocks and / or frames and metadata from the encoded audio signal, for example, through operations such as decompression, deformatting, and demultiplexing ("DEMUX"). The extracted audio data blocks and / or frames may be decoded into audio data elements or samples by a decoder unit ("NGC decoder"). The decoder (100) is further configured to determine a profile for a particular playback environment in which the audio content will be rendered and to select a dynamic range compression curve from the metadata extracted from the encoded audio signal. The digital audio processing unit ("DAP") is configured to apply DRC or other operations to the audio data elements or samples to generate an audio signal that drives the audio channels in the particular playback environment. The decoder (100) can calculate and apply a DRC gain and a selected dynamic range compression curve based on the audio data block or frame. The decoder (100) can also adjust the output dialogue loudness level based on a reference dialogue loudness level associated with the selected dynamic range compression curve and the dialogue loudness level in the metadata extracted from the encoded audio signal. The decoder (100) can then apply a gain limiter specific to the audio content and the playback scenario associated with the particular playback environment. In this way, the decoder (100) can render / play back the audio content in a manner tailored to the playback scenario.

[0154] FIG. 5A illustrates another exemplary decoder (which may be the same as decoder 100 of FIG. 5). As shown in FIG. 5A, the decoder of FIG. 5A is configured to extract audio data blocks and / or frames and metadata from an encoded audio signal, e.g., through operations such as decompression, deformatting, and demultiplexing ("DEMUX"). The extracted audio data blocks and / or frames may be decoded by a decoder unit ("Decode") into audio data elements or samples. The decoder of FIG. 5A is further configured to perform DRC gain calculations based on a default compression curve, a smoothing constant associated with the default compression curve, etc., for a set of default gains. The decoder of FIG. 5A is further configured to extract a set of differential gains for non-default gain profiles from profile-related metadata in the metadata, determine a set of non-differential gains for the non-default gain profiles in the decoder of FIG. 5A where the audio content is to be rendered, and apply the set of non-differential gains and other operations to the audio data elements or samples to generate a DRC-enhanced audio output that drives the audio channels in a particular playback environment. The decoder of Figure 5A can render / play audio content according to a non-default gain profile whether or not the decoder of Figure 5A itself implements support for performing a set of gain generation operations to obtain a set of non-differential gains directly for the non-default gain profile.

[0155] 6A-6D illustrate an exemplary process flow. In some embodiments, one or more computing devices or units in a media processing system may perform this process flow.

[0156] 6A illustrates an exemplary process flow that may be implemented by an audio decoder described herein. In block 602 of FIG. 6A, a first device (such as, for example, audio decoder 100 of FIG. 1A) receives an audio signal that includes audio content and definition data for one or more dynamic range compression curves.

[0157] In block 604, the first device determines the particular playback environment.

[0158] In block 606, the first device establishes a particular dynamic range compression curve for the particular playback environment based on the definition data for the one or more dynamic range compression curves extracted from the audio signal.

[0159] In block 608, the first device performs one or more dynamic range control (DRC) operations on one or more portions of audio content extracted from the audio signal, the one or more DRC operations based at least in part on one or more DRC gains derived from a particular dynamic range compression curve.

[0160] In one embodiment, the definition data for the one or more dynamic range compression curves includes an attack time, a release time, or a reference loudness level associated with at least one of the one or more dynamic range compression curves.

[0161] In one embodiment, the first device is further configured to perform the steps of: calculating one or more loudness levels for the one or more portions of audio content; determining the one or more DRC gains based on the particular dynamic range compression curve and the one or more loudness levels for the one or more portions of the audio content;

[0162] In some embodiments, at least one of the loudness levels calculated for the one or more portions of the audio content is one or more of: a specific loudness level related to one or more frequency bands, a broadband loudness level across a broadband range, a wideband loudness level across a wideband range, a broadband loudness level across multiple frequency bands, a wideband loudness level across multiple frequency bands, etc.

[0163] In some embodiments, at least one of the loudness levels calculated for the one or more portions of the audio content is one or more of an instantaneous loudness level or a loudness level smoothed over one or more time intervals.

[0164] In one embodiment, the one or more operations relate to one or more of adjusting a dialogue loudness level, gain smoothing, gain limiting, dynamic equalization, noise compensation, and the like.

[0165] In one embodiment, the first device is further configured to: extract one or more dialogue loudness levels from the encoded audio signal; adjust the one or more dialogue loudness levels to one or more reference dialogue loudness levels; etc.

[0166] In one embodiment, the first device is further configured to: extract one or more auditory scene analysis (ASA) parameters from the encoded audio signal; vary one or more time constants used in smoothing a gain applied to the audio content, the gain related to one or more of the one or more DRC gains; perform gain smoothing or gain limiting, etc.

[0167] In one embodiment, the first device is further configured to: determine that a reset event occurs in the one or more portions of the audio content based on an indication of a reset event, the indication of the reset being extracted from the encoded audio signal; and in response to determining that the reset event occurs in the one or more portions of the audio content, perform one or more actions on one or more gain smoothing operations that are running at the time of determining that the reset event occurs in the one or more portions of the audio content.

[0168] In one embodiment, the first device is further configured to: maintain a histogram of instantaneous loudness levels, the histogram being populated with instantaneous loudness levels calculated from a time interval in the audio content; determine whether a specific loudness level is above a threshold in a high probability region of the histogram, the specific loudness level being calculated from a portion of the audio content; and in response to determining that the specific loudness level is above the threshold in the high probability region of the histogram, determine that a loudness transition is occurring and, for example, shorten a time constant used in gain smoothing to speed up the loudness transition.

[0169] 6B illustrates an exemplary process flow that may be implemented by an audio encoder described herein. In block 652 of FIG. 6B, a second device (such as audio encoder 150 of FIG. 1B) receives audio content in a source audio format.

[0170] In block 654, the second device obtains definition data for one or more dynamic range compression curves.

[0171] In block 656, the second device generates an audio signal including the audio content and the definition data for the one or more dynamic range compression curves.

[0172] In one embodiment, the second device is further configured to: determine one or more identifiers for the one or more dynamic range compression curves; and retrieve the definition data for the one or more dynamic range compression curves from a reference data repository based on the one or more identifiers.

[0173] In one embodiment, the second device is further configured to: calculate one or more dialogue loudness levels for the one or more portions of the audio content; and encode the one or more dialogue loudness levels together with the one or more portions of the audio content into the encoded audio signal.

[0174] In one embodiment, the second device is configured to: perform an auditory event scene (ASA) on the one or more portions of the audio content; generate one or more ASA parameters based on the results of the ASA on the one or more portions of the audio content; encode the one or more ASA parameters together with the one or more portions of the audio content into the encoded audio signal; and the like.

[0175] In one embodiment, the second device is further configured to: determine that one or more reset events occur in the one or more portions of the audio content; and encode one or more indicators of the one or more reset events together with the one or more portions of the audio content into the encoded audio signal.

[0176] In one embodiment, the second device is further configured to encode the one or more portions of the audio content into one or more audio data frames or audio data blocks.

[0177] In one embodiment, a first DRC gain of the one or more DRC gains applies to each channel in a first proper subset of the set of all channels in a particular speaker configuration corresponding to that particular playback environment, while a second, different DRC gain of the one or more DRC gains applies to each channel in a second proper subset of the set of all channels in the particular speaker configuration corresponding to that particular playback environment.

[0178] In one embodiment, a first DRC gain of the one or more DRC gains applies to a first frequency band and a second, different DRC gain of the one or more DRC gains applies to a second, different frequency band.

[0179] In some embodiments, the one or more portions of the audio content include one or more audio data frames or audio data blocks. In some embodiments, the encoded audio signal is part of an audiovisual signal.

[0180] In one embodiment, the one or more DRC gains are defined in the loudness domain.

[0181]

[0033] Figure 6C illustrates an exemplary process flow that may be implemented by an audio decoder described herein. At block 662 of Figure 6C, a third device (e.g., audio decoder 100 of Figure 1A, the audio decoder of Figure 5, the audio decoder of Figure 5A, etc.) receives an audio signal that includes audio content, definition data for one or more dynamic range compression (DRC) curves, and one or more sets of differential gains.

[0182] In block 664, the third device identifies, from among the one or more sets of differential gains, a specific set of differential gains for a gain profile in a specific playback environment, and also identifies, from the one or more DRC curves, a default DRC curve associated with the specific set of differential gains.

[0183] In block 666, the third device generates a set of default gains based at least in part on the default DRC curve.

[0184] At block 668, based at least in part on the combination of the set of default gains and the particular set of differential gains, the third device performs one or more operations on one or more portions of the audio content extracted from the audio signal.

[0185] In one embodiment, the set of default gains includes non-differential gains generated by performing a set of gain generation operations based at least in part on the default DRC curve.

[0186] In some embodiments, the default DRC curve represents a default gain profile. In some embodiments, the particular set of differential gains relative to the default DRC curve represents a non-default gain profile. In some embodiments, the audio signal does not include definition data for a non-default DRC curve corresponding to the non-default gain profile.

[0187] In some embodiments, the particular set of differential gains comprises a gain difference between a set of non-differential non-default gains generated for a non-default gain profile and a set of non-differential default gains generated for the default gain profile represented by the default DRC curve, and the set of non-differential non-default gains and the set of non-differential default gains may be generated by an upstream audio decoder that encodes the audio signal.

[0188] In some embodiments, at least one of the set of non-differential non-default gains or the set of non-differential default gains is not provided as part of the audio signal.

[0189] 6D illustrates an exemplary process flow that may be implemented by an audio decoder described herein. In block 672 of FIG. 6D, a fourth device (e.g., audio encoder 150 of FIG. 1A, the audio encoder of FIG. 5, etc.) receives audio content in a source audio format.

[0190] In block 674, the fourth device generates a set of default gains based, at least in part, on a default dynamic range compression (DRC) curve representing a default gain profile.

[0191] In block 676, the fourth device generates a set of non-default gains for the non-default gain profile.

[0192] At block 678, based at least in part on the set of default gains and the set of non-default gains, a fourth device generates a set of differential gains, the set of differential gains representing the non-default gain profile relative to the default DRC curve.

[0193] At block 680, a fourth device generates an audio signal including the audio content and the definition data for one or more DRC curves and for one or more sets of differential gains, the one or more sets of differential gains including the set of differential gains.

[0194] In some embodiments, the non-default gain profile is represented by a DRC curve. In some embodiments, the audio signal does not include definition data for the DRC curve representing the non-default gain profile. In some embodiments, the non-default gain profile is not represented by a DRC curve.

[0195] In one embodiment, an apparatus having a processor and configured to perform any of the methods described herein.

[0196] In one embodiment, a non-transitory computer-readable storage medium containing software instructions that, when executed by one or more processors, cause the performance of any of the methods described herein. It is noted that although separate embodiments are discussed herein, any combination of the embodiments and / or sub-embodiments discussed herein may be combined to form further embodiments.

[0197] 19. Implementation Mechanism - Hardware Overview According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-configured to perform the techniques, or may include digital electronic devices, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), that are persistently programmed to perform the techniques. Alternatively, the special-purpose computing devices may include one or more general-purpose hardware processors that are programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such special-purpose computing devices may combine custom hard-configured logic, ASICs, or FPGAs with custom programming to achieve the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices incorporating hard-configured and / or program logic to implement the techniques.

[0198] 7 is a block diagram illustrating a computer system 700 in which an embodiment of the present invention may be implemented. Computer system 700 includes a bus 702 or other communication mechanism for communicating information, and a hardware processor 704 coupled to bus 702 for processing information. Hardware processor 704 may be, for example, a general-purpose microprocessor.

[0199] Computer system 700 also includes a main memory 706, coupled to bus 702, such as a random access memory (RAM) or other dynamic storage device, for storing information and instructions to be executed by processor 704. Main memory 706 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 704. Such instructions, when stored in a non-transitory storage medium accessible to processor 704, cause computer system 700 to become a device-specific, special-purpose machine for performing operations specified in the instructions.

[0200] Computer system 700 further includes a read-only memory (ROM) 708 or other static storage device coupled to bus 702 for storing static information and instructions for processor 704. A storage device 710, such as a magnetic disk or optical disk, is provided and coupled to bus 702 for storing information and instructions.

[0201] Computer system 700 may be coupled via bus 702 to a display 712, such as a liquid crystal display (LCD), for displaying information to a computer user. An input device 714, including alphanumeric and other keys, is coupled to bus 702 for communicating information and command selections to processor 704. Another type of user input device is a cursor control 716, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 704 and for controlling cursor movement on display 712. This input device typically has two degrees of freedom along two axes, a first axis (e.g., x) and a second axis (e.g., y), allowing the device to specify a position in a plane.

[0202] Computer system 700 may implement the techniques described herein using device-specific hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic that, in combination with the computer system, configures or programs computer system 700 as a special-purpose machine. According to one embodiment, the techniques described herein are performed by computer system 700 in response to processor 704 executing one or more sequences of one or more instructions contained in main memory 706. Such instructions may be read into main memory 706 from another storage medium, such as storage device 710. Execution of the sequences of instructions contained in main memory 706 causes processor 704 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0203] The term "storage medium" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 710. Volatile media include dynamic memory, such as main memory 706. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs, NVRAM, and any other memory chip or cartridge.

[0204] Storage media is distinct from, but may be used in conjunction with, transmission media. Transmission media participates in transferring information between storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 702. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0205] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 704 for execution. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 700 can receive the data on the telephone line and convert the data to an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on bus 702. Bus 702 carries the data to main memory 706, from which processor 704 retrieves and executes the instructions. The instructions received by main memory 706 may optionally be stored on storage device 710 either before or after execution by processor 704.

[0206] Computer system 700 also includes a communication interface 718 coupled to bus 702. The communication interface 718 provides a two-way data communication coupling to a network link 720 that is connected to a local network 722. For example, communication interface 718 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a corresponding type of modem to provide a data communication connection to a telephone line. As another example, communication interface 718 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 718 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0207] Network link 720 typically provides data communication through one or more networks to other data devices. For example, network link 720 may provide a connection through local network 722 to a host computer 724 or to data equipment operated by an Internet Service Provider (ISP) 726. ISP 726 provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 728. Local network 722 and the Internet 728 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 720 and through communication interface 718, which carry the digital data to and from computer system 700, are exemplary forms of transmission media.

[0208] Computer system 700 can send messages and receive data, including program code, through the network(s), network link 720 and communication interface 718. In the Internet example, a server 730 might transmit a requested code for an application program through Internet 728, ISP 726, local network 722 and communication interface 718.

[0209] The received code may be executed by processor 704 as it is received, and / or stored in storage device 710, or other non-volatile storage for later execution.

[0210] 20. Equivalents, Extensions, Substitutions, etc. In the foregoing specification, exemplary embodiments of the invention have been described, with reference to numerous specific details that may vary depending on the implementation. As such, the sole and exclusive indication of what is, and what is intended by applicant to be, the invention is the claims of any patent issued on this application, in the specific form in which such claims are patented, including any subsequent amendments thereto. The definitions, if any, expressly set forth herein for terms contained in such claims will govern the meaning of such terms as used in the claims. Accordingly, no limitation, element, attribute, feature, advantage, or characteristic not expressly recited in a claim should in any way limit the scope of such claim. Accordingly, the specification and drawings are to be regarded in an illustrative and not a restrictive sense.

[0211] Several aspects will be described. [Aspect 1] receiving an audio signal including audio content and one or more sets of differential gains; identifying, from the one or more sets of differential gains, a particular set of differential gains for a gain profile in a particular playback environment; generating a set of default gains based at least on a default dynamic range compression (DRC) curve associated with the particular set of differential gains; performing one or more operations on one or more portions of the audio content extracted from the audio signal based at least in part on a combination of the set of default gains and the specified set of differential gains. A method implemented by one or more computers. [Aspect 2] 2. The method of aspect 1, wherein the set of default gains comprises non-differential gains generated by performing a set of gain generation operations based at least in part on the default DRC curve. Aspect 3 3. The method of claim 1 or 2, wherein the default DRC curve represents a default gain profile. Aspect 4 Aspect 4. The method of any one of aspects 1-3, wherein the particular set of differential gains relative to the default DRC curve represents a non-default gain profile. Aspect 5 5. The method of claim 4, wherein the audio signal does not include definition data for a non-default DRC curve corresponding to the non-default gain profile. Aspect 6 6. The method of any one of aspects 1 to 5, wherein the particular set of differential gains comprises a gain difference between a set of non-differential non-default gains generated for a non-default gain profile and a set of non-differential default gains generated for the default gain profile represented by the default DRC curve. Aspect 7 7. The method of claim 6, wherein the set of non-differential non-default gains and the set of non-differential default gains are generated by an upstream audio decoder that encodes the audio signal. Aspect 8 7. The method of embodiment 6, wherein at least one of the set of non-differential non-default gains or the set of non-differential default gains is not provided as part of the audio signal. Aspect 9 9. The method of any one of aspects 1 to 8, wherein the definition data for the one or more DRC curves includes one or more of an attack time, a release time, or a reference loudness level associated with at least one of the one or more DRC curves. Aspect 10 10. The method of claim 9, wherein the reference loudness level represents a target range of playback levels for rendering the audio content by an audio decoder. Aspect 11 calculating one or more loudness levels for the one or more portions of the audio content; generating a set of non-differential non-default gains based on the set of non-differential default gains and the specified set of differential gains; and applying the set of non-differential non-default gains to the one or more portions of the audio content. 11. The method of any one of embodiments 1 to 10. Aspect 12 12. The method of claim 11, wherein at least one of the one or more loudness levels calculated for the one or more portions of the audio content is one or more of a specific loudness level for one or more frequency bands, a broadband loudness level across a broadband range, a wideband loudness level across a wideband range, a broadband loudness level across multiple frequency ranges, or a wideband loudness level across multiple frequency ranges. Aspect 13 12. The method of claim 11, wherein at least one of the one or more loudness levels calculated for the one or more portions of the audio content is one or more of an instantaneous loudness level or a loudness level smoothed over one or more time intervals. Aspect 14 14. The method of any one of aspects 1 to 13, wherein the one or more operations include one or more operations related to one or more of adjusting a dialogue loudness level, gain smoothing, gain limiting, dynamic equalization, or noise compensation. Aspect 15 15. The method of any one of aspects 1 to 14, wherein the method is performed by an audio decoding device, and the default DRC curve is defined in the audio decoding device. Aspect 16 receiving definition data for one or more dynamic range compression (DRC) curves; and identifying, among the one or more DRC curves, a default DRC curve associated with the particular set of differential gains. 16. The method of any one of embodiments 1 to 15. Aspect 17 extracting one or more auditory scene analysis (ASA) parameters from the encoded audio signal; and varying one or more time constants used in smoothing the gain applied to the audio content. 17. The method of any one of embodiments 1 to 16. Aspect 18 determining that a reset event occurs in the one or more portions of the audio content based on an indicator of a reset event, the indicator of the reset being extracted from the encoded audio signal; and in response to determining that the reset event occurs in the one or more portions of the audio content, performing one or more actions on one or more gain smoothing operations that are running at the time of determining that the reset event occurs in the one or more portions of the audio content. 18. The method of any one of embodiments 1 to 17. Aspect 19 20. The method of claim 18, wherein at least one of the one or more smoothing operations uses a first smoothing time constant before the reset event, and wherein the at least one of the one or more smoothing operations uses a second smoothing time constant less than the first smoothing time constant in response to determining that the reset event occurs. Aspect 20 maintaining a histogram of instantaneous loudness levels, the histogram being populated with instantaneous loudness levels calculated from a time interval in the audio content; determining whether a specific loudness level is below a threshold in a high probability region of the histogram, the specific loudness level being calculated from a portion of the audio content; In response to determining that the specific loudness level is below the threshold in the high probability region of the histogram: determining that a loudness transition is occurring; shortening a time constant used in gain smoothing to speed up the loudness transition. 20. The method of any one of embodiments 1 to 19. Aspect 21 21. The method of any one of aspects 1 to 20, wherein the specific set of differential gains includes a first differential gain associated with each channel in a first proper subset of the set of all channels in a particular speaker configuration, and the specific set of differential gains includes a second differential gain associated with each channel in a second proper subset of the set of all channels in the particular speaker configuration. Aspect 22 22. The method of any one of aspects 1-21, wherein the specific set of differential gains includes a first differential gain associated with a first frequency band, and the specific set of differential gains includes a second, different differential gain associated with a second, different frequency band. Aspect 23 23. The method of any one of aspects 1 to 22, wherein the one or more portions of the audio content include one or more of an audio data frame, an audio data block, or an audio sample. Aspect 24 24. The method of any one of aspects 1-23, wherein the particular set of differential gains is defined in the loudness domain. Aspect 25 25. The method of any one of aspects 1-24, wherein the encoded audio signal is part of an audiovisual signal. Aspect 26 receiving audio content in a source audio format; generating a set of default gains based at least in part on a default dynamic range compression (DRC) curve, the default DRC curve representing a default gain profile; generating a set of non-default payoffs for the non-default payoff profile; generating the set of differential gains based at least in part on the set of default gains and the set of non-default gains, the set of differential gains representing the non-default gain profile relative to the default DRC curve; generating an audio signal comprising the audio content and the one or more sets of differential gains comprising the set of differential gains; A method performed by one or more computing devices. Aspect 27 27. The method of embodiment 26, wherein the non-default gain profile is represented by a DRC curve. Aspect 28 28. The method of claim 27, wherein the audio signal does not include definition data for the DRC curve representing the non-default gain profile. Aspect 29 29. The method of any one of embodiments 26-28, wherein the non-default gain profile is not represented by a DRC curve. Aspect 30 determining one or more identifiers for the one or more dynamic range compression curves; and retrieving the definition data for the one or more dynamic range compression curves from a reference data store based on the one or more identifiers. 30. The method of any one of embodiments 26 to 29. Aspect 31 31. The method of any one of aspects 26 to 30, wherein the set of default gains includes first non-differential gains generated by performing a first set of gain generation operations based at least in part on the default DRC curve, and the set of non-default gains includes second non-differential gains generated by performing a second set of gain generation operations for the non-default gain profile. Aspect 32 calculating one or more dialogue loudness levels for one or more portions of the audio content; encoding the one or more dialogue loudness levels together with the one or more portions of the audio content into the encoded audio signal. 32. The method of any one of embodiments 26 to 31. Aspect 33 33. The method of embodiment 32, wherein at least one of the one or more dialogue loudness levels is determined from one or more audio tracks containing dialogue audio content. Aspect 34 performing an auditory scene analysis (ASA) on the one or more portions of the audio content; generating one or more ASA parameters based on the ASA results for the one or more portions of the audio content; encoding the one or more ASA parameters together with the one or more portions of the audio content into the encoded audio signal. 34. The method of any one of embodiments 26 to 33. Aspect 35 determining that one or more reset events occur in one or more portions of the audio content; encoding one or more indications of the one or more reset events into the encoded audio signal along with the one or more portions of the audio content. 35. The method of any one of embodiments 26 to 34. Aspect 36 36. The method of any one of aspects 26 to 35, further comprising encoding one or more portions of the audio content into one or more audio data frames or audio data blocks. Aspect 37 37. The method of any one of embodiments 26 to 36, wherein at least one of the one or more dynamic range compression curves is defined in the loudness domain. Aspect 38 38. The method of any one of aspects 26-37, wherein the encoded audio signal is part of an audiovisual signal. Aspect 39 39. The method of any one of aspects 26 to 38, wherein the definition data for the one or more dynamic range compression curves includes one or more sets of parameters, and at least one set in the one or more sets of parameters represents one or more of a lookup table, a curve, or a multi-segment piecewise straight line. Aspect 40 40. The method of any one of embodiments 26-39, wherein the encoded audio signal includes an indicator for selecting a DRC curve defined in a receiving device as the default DRC curve. Aspect 41 sending definition data for DRC curves in the encoded audio signal; and including an indicator for selecting the default DRC curve from the one or more DRC curves. 41. The method of any one of embodiments 26 to 40. Aspect 42 42. A media processing system configured to perform the method of any one of aspects 1 to 41. Aspect 43 42. An apparatus having a processor configured to perform the method of any one of aspects 1 to 41. Aspect 44 42. A non-transitory computer-readable storage medium comprising software instructions that, when executed by one or more processors, cause performance of the method of any one of aspects 1 to 41.

Claims

1. 1. A method for dynamic range control (DRC) of an audio signal, the method comprising: receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, the audio signal including audio sample data for each channel of the reference channel configuration and encoder-generated DRC metadata, the encoder-generated DRC metadata including DRC gains for multiple channel configurations, the DRC gains for the multiple channel configurations including a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the audio sample data to generate output audio sample data for each channel of the playback channel configuration; the playback channel configuration is a two-channel configuration; method.

2. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, perform the following: receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, the audio signal including audio sample data for each channel of the reference channel configuration and encoder-generated dynamic range control (DRC) metadata, the encoder-generated DRC metadata including DRC gains for multiple channel configurations, the DRC gains for the multiple channel configurations including a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the audio sample data to generate output audio sample data for each channel of the playback channel configuration; the playback channel configuration is a two-channel configuration; A non-transitory computer-readable storage medium.

3. 1. An audio signal processing apparatus for dynamic range control of an audio signal, the audio signal processing apparatus comprising: receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, the audio signal including audio sample data for each channel of the reference channel configuration and encoder-generated DRC metadata, the encoder-generated DRC metadata including DRC gains for multiple channel configurations, the DRC gains for the multiple channel configurations including a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the audio sample data to generate output audio sample data for each channel of the playback channel configuration; the playback channel configuration is a two-channel configuration; Device.

4. A computer program product having executable instructions for performing the method of claim 1 when executed on a computer.