Dynamic Range Control for Diverse Reproduction Environments

By transmitting dynamic range compression curves and differential gains, the audio encoder allows decoders to customize processing for varied playback environments, addressing inconsistent loudness and intelligibility issues in media devices.

JP7717925B2Active Publication Date: 2025-08-04DOLBY LABORATORIES LICENSING CORP +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024139393
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-02-10
Filing Date
2024-08-21
Publication Date
2025-08-04
Estimated Expiration
2034-09-08

AI Technical Summary

Technical Problem

Media processing devices struggle to maintain consistent loudness and intelligibility across various media formats and playback environments, particularly with high-quality, wide-bandwidth audio content, due to irreversible audio processing assumptions made by encoders.

Method used

The audio encoder transmits dynamic range compression curves and differential gains, allowing decoders to customize audio processing based on playback environments, supporting flexible gain profiles and maintaining perceptual quality.

Benefits of technology

This approach enables consistent loudness and intelligibility across diverse playback scenarios, supporting new gain profiles and environments without locking into irreversible processing, ensuring optimal audio reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717925000006
    Figure 0007717925000006
  • Figure 0007717925000007
    Figure 0007717925000007
  • Figure 0007717925000008
    Figure 0007717925000008
Patent Text Reader

Abstract

To provide dynamic range control for various reproduction environment.SOLUTION: In an audio encoder, default gain is generated based on a default dynamic range compression (DRC) curve, for an audio content received in source audio format, and a non-default gain is generated for non-default gain profile. A finite difference gain is generated based on the default gain and the non-default gain. An audio signal including an audio content, a default DRC curve and the finite difference gain is generated. In the audio decoder, the default DRC curve and the finite difference gain are identified from the audio signal. A default gain is generated again on the basis of the default DRC curve. Based on the combination of the regenerated default gain and the finite difference gain, operation is executed for the audio content extracted from the audio signal.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 61 / 877,230, filed on September 12, 2013; U.S. Provisional Patent Application No. 61 / 891,324, filed on October 15, 2013; and U.S. Provisional Patent Application No. 61 / 938,043, filed on February 10, 2014. The content of each application is hereby incorporated by reference in its entirety.

[0002] Technique The present invention generally relates to techniques that can be used to apply dynamic range control and other types of audio processing operations to audio signals in any of a wide variety of playback environments, more particularly to the processing of audio signals.

Background Art

[0003] The growing popularity of media consumption devices has created new opportunities and challenges for the creators and distributors of media content for playback on such devices, or for the designers and manufacturers of such devices. Many consumer devices can play a wide variety of media content types and formats, including those often associated with high - quality, wide - bandwidth, and wide dynamic range audio content for HDTV, Blu - ray, or DVD. Media processing devices can be used to play this type of audio content on their internal acoustic transducers or on external transducers such as headphones. However, media processing devices generally cannot play this content with consistent loudness and intelligibility across a variety of media formats and content types.

[0004] The approaches described in this section could have been pursued, but are not necessarily approaches that were previously conceived or pursued. Therefore, unless otherwise noted, none of the approaches described in this section should be assumed to be prior art merely by virtue of being included in this section. Similarly, unless otherwise noted, problems identified with respect to one or more approaches should not be assumed to have been recognized in any prior art based on this section.

Brief Description of the Drawings

[0005] The present invention is shown, by way of example and not limitation, in the accompanying drawings. In the drawings, like reference numerals refer to like elements.

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 5A

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 7

Figure 7

[0006] Exemplary embodiments are described in this document for applying dynamic range control and other types of audio processing operations to an audio signal in any of a wide variety of playback environments. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. On the other hand, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, obfuscating, or burying the present invention.

[0007] The exemplary embodiments are described in this document according to the following outline. 1. General Overview 2. Dynamic Range Control 3. Audio Decoder 4. Audio Encoder 5. Dynamic Range Compression Curve 6. DRC Gain, Gain Limiting, and Gain Smoothing 7. Input Smoothing and Gain Smoothing 8. DRC Over Multiple Frequency Bands 9. Volume Adjustment in Loudness Region 10. Gain Profile by Differential Gain 11. Additional Operations Related to Gain 12. Specific and Broadband (or Wideband) Loudness Levels 13. Individual Gains for Individual Subsets of Channels 14. Auditory Scene Analysis 15. Loudness Level Transition 16. Reset 17. Gain Provided by Encoder 18. Exemplary Systems and Process Flows 19. Implementation Mechanisms - Overview of Hardware 20. Equivalents, Extensions, Alternatives, etc.

[0008] 〈1. General Overview〉 This overview presents a basic description of some aspects of the embodiments of the present invention. It should be noted that this overview is not an inclusive or comprehensive summary of the aspects of the embodiments. Furthermore, it is not intended that this overview be understood as identifying any particularly significant aspect or element of the embodiments, nor generally as delimiting any scope of the present invention, particularly of the embodiments. This overview simply presents some concepts related to its exemplary embodiments in a condensed and simplified form and should be understood merely as a conceptual introduction to the more detailed description of the subsequent exemplary embodiments. It should be noted that although separate embodiments are discussed in this document, any combination of the embodiments and / or partial embodiments discussed in this document may be combined to form further embodiments.

[0009] In some approaches, the encoder assumes that the audio content is encoded for a particular environment for the purpose of dynamic range control, and for that particular environment, determines audio processing parameters such as gains for dynamic range control and the like. The gains determined by the encoder under these approaches are typically smoothed with some time constant (e.g., in an exponential decay function, etc.) over some time interval, etc. Further, the gains determined by the encoder under these approaches may be incorporated with gain limitations that ensure that the signal does not exceed the clipping level for the assumed environment. Thus, the gains encoded into the audio signal along with the audio information by the encoder under these approaches are the result of many different effects and are irreversible. A decoder receiving the gain under these approaches will not be able to distinguish which part of the gain is for dynamic range control, which part of the gain is for gain smoothing, which part of the gain is for gain limitation, etc.

[0010] Under the techniques described in this document, the audio encoder does not assume that only a particular playback environment in the audio decoder needs to be supported. In one embodiment, the audio encoder transmits an encoded audio signal having audio content from which a correct loudness level (e.g., without clipping, etc.) can be determined. The audio encoder may also transmit one or more dynamic range compression curves to the audio decoder. Any of the one or more dynamic range compression curves may be standard-based, proprietary, customized, content-provider specific, etc. A reference loudness level, attack time, release time, etc. may be transmitted by the audio encoder as part of or in relation to the one or more dynamic range compression curves.

[0011] In some embodiments, the audio encoder implements auditory scene analysis (ASA) techniques, uses the ASA techniques to detect auditory events in the audio content, and sends one or more ASA parameters describing the detected auditory events to an audio decoder.

[0012] In some embodiments, the audio encoder can also be configured to detect a reset event in the audio content and send an indication of the reset event to a downstream device, such as an audio decoder, in a time-synchronized manner along with the audio content.

[0013] In some embodiments, the audio encoder may be configured to calculate one or more sets of gains (such as DRC gains etc.) for individual portions of the audio content (such as audio data blocks, audio data frames, etc.), and encode the set of gains, together with the individual portions of the audio content, into the encoded audio signal. In some embodiments, the set of gains generated by the audio encoder corresponds to one or more gain profiles (such as those shown in Table 1). In some embodiments, Huffman coding, differential coding, etc. may be used to encode the set of gains into components, subdivisions, etc. of the audio data frames, or to read the set of gains from the components, subdivisions, etc. These components, subdivisions, etc. may be referred to as sub-frames of the audio data frame. Different sets of gains may correspond to different sets of sub-frames. Each set of gains or each set of sub-frames may have two or more temporal components (such as sub-frames, etc.). In some embodiments, the bitstream formatter in the audio encoder described herein may use one or more for loops to write one or more sets of gains, as differential data codes, together into one or more sets of sub-frames in the audio data frame. Correspondingly, the bitstream parser in the audio decoder described herein may read any of the one or more sets of gains encoded as the differential data code from the one or more sets of sub-frames in the audio data frame. [[ID=!]]

[0014] In some embodiments, the audio encoder determines the dialog loudness level in the audio content to be encoded into the encoded audio signal, and sends the dialog loudness level to the audio decoder together with the audio content.

[0015] In some embodiments, the audio encoder sends a default dynamic compression curve for a default gain profile in the playback environment or scenario to a downstream receiving audio decoder. In some embodiments, the audio encoder assumes that the downstream receiving audio decoder will use the default dynamic compression curve for the default gain profile in the playback environment or scenario. In some embodiments, the audio encoder sends an indication to the downstream receiving audio decoder as to which of one or more dynamic compression curves defined in the downstream receiving audio decoder should be used in the playback environment or scenario. In some embodiments, for each of one or more non-default gain profiles, the audio encoder sends, as part of the metadata carried by the encoded audio signal, a dynamic compression curve corresponding to that non-default profile (e.g., non-default, etc.). The techniques described herein allow multiple sets of differential gains related to the default compression curve to be generated by an upstream encoder and sent to a downstream decoder. This allows for a relatively low bitrate to be maintained compared to transmitting full gain values while allowing a great deal of freedom in the design of the DRC compressor in the decoder (e.g., the process of calculating gain based on a compression curve and smoothing operations, etc.). Merely by way of example, the default profile or default DRC curve has been referred to as one for which differential gains for non-default profiles or non-default DRC curves can be specifically calculated in relation to it. However, this is merely for illustration, and there is no strict need to distinguish between default and non-default profiles (e.g., in a media data stream, etc.). In various embodiments, this is because all other profiles can be differential gains compared to the same specific (e.g., "default", etc.) compression curve. In the usage herein, "gain profile" may be referred to as the DRC mode as the operating mode of a compressor that performs DRC operations.In some embodiments, the DRC mode is related to the specific type of the playback device (such as AVR, TV, or tablet) and / or the environment (noisy, quiet, or late at night). Each DRC mode can be associated with a gain profile. The gain profile may be represented by definition data, and based on the definition data, the compressor performs DRC operations. In some embodiments, the gain profile can be a DRC curve (possibly parameterized) and a time constant used in the DRC operation. In some embodiments, the gain profile can be a set of DRC gains as the output of the DRC operation in response to the audio signal. The profiles of different DRC modes may correspond to different amounts of compression.

[0016] In some embodiments, the audio encoder determines a set of default (e.g., full DRC and non-DRC, full DRC, etc.) gains for the audio content based on a default dynamic range compression curve corresponding to the default gain profile, and for each of one or more non-default gain profiles, determines a set of non-default (e.g., full DRC and non-DRC, full DRC, etc.) gains for the same audio content. Then, the audio encoder can determine the gain difference between the set of default (e.g., full DRC and non-DRC, full DRC, etc.) gains for the default gain profile and the set of non-default (e.g., full DRC and non-DRC, full DRC, etc.) gains for the non-default gain profiles, and include the gain difference in a set of differential gains. Instead of sending a dynamic compression curve (such as non-default, etc.) for a non-default playback environment or scenario related non-default profile, the audio encoder can send the set of differential gains as part of the metadata carried by the encoded audio signal, instead of or in addition to the non-default dynamic compression curve.

[0017] The set of differential gains may be smaller in size than the set of non-default (e.g., full DRC and non-DRC, full DRC, etc.) gains. Thus, transmitting differential gains instead of non-differential (e.g., full DRC and non-DRC, full DRC, etc.) gains may sometimes require a lower bit rate compared to directly transmitting non-differential (e.g., full DRC and non-DRC, full DRC, etc.) gains.

[0018] The audio decoders that receive the encoded audio signals described in this document may be provided by different manufacturers and implemented with different components and designs. The audio decoders may be released to end users at different times, or may be updated with different versions of hardware, software, and firmware. As a result, those audio decoders may have different audio processing functions. In some embodiments, a number of audio decoders may be equipped with a function to support a limited set of gain profiles, such as a default gain profile defined by a standard, proprietary requirements, etc. A number of audio decoders may be configured with a function to perform related gain generation operations to generate gains for the default gain profile based on a default dynamic range compression curve representing the default gain profile. Transmitting the default dynamic range compression curve for the default gain profile in the audio signal may be more efficient than transmitting the gains generated / calculated for the default gain profile in the audio signal.

[0019] For a non-default gain profile, the audio encoder can pre-generate a differential gain by referring to a specific default dynamic range compression curve corresponding to a specific default gain profile. In response to receiving the differential gain in the audio signal generated by the audio encoder, the audio decoder generates a default gain based on the default dynamic range compression curve received in the audio signal, combines the received differential gain and the generated default gain into a non-default gain for the non-default gain profile, and can render the received audio content while applying the non-default gain to the audio content decoded from the audio signal. In some embodiments, the non-default gain profile may be used to ensure the limitations of the default dynamic range compression curve.

[0020] The techniques described herein can be used to provide flexible support for new gain profiles, features, or improvements. In some embodiments, at least one gain profile, whether default or non-default, cannot be easily represented using a dynamic range compression curve. In some embodiments, at least one gain profile may be specific to a particular audio content (such as a particular movie, etc.). It may also be necessary to transmit more parameters, smoothing constants, etc., that can be carried in the encoded audio signal, in the encoded audio signal, for the representation of the non-default gain profile (such as a parameterized DRC curve). In some embodiments, at least one gain profile may be specific to a particular audio content provider (such as a particular studio, etc.).

[0021] As such, the audio encoder described in this document can be leading in supporting a new gain profile. It does so by implementing a gain generation operation for the new gain profile and a gain generation operation for the default gain profile related to the new gain profile. The downstream receiving audio decoder does not need to perform a gain generation operation for the new gain profile. Rather, the audio decoder can support the new gain profile by leveraging the non-default differential gain generated by the audio encoder without the audio decoder performing a gain generation operation for the new gain profile.

[0022] In some embodiments, in the profile-related metadata encoded in the encoded audio signal, one or more sets of dynamic range compression curves (such as default, etc.) and one or more sets of differential gains (such as non-default, etc.) are structured, indexed, etc. according to their respective gain profiles. In some embodiments, the relationship between the set of non-default differential gains and the default dynamic range compression curve may be indicated in the profile-related metadata. This can be particularly useful when there are two or more default dynamic range compression curves in the metadata or when they are not in the metadata but are defined in the downstream decoder. Based on the relationship indicated in the profile-related metadata, the receiving audio decoder can determine which default dynamic range compression curve should be used to generate the set of default gains. The generated gains can then be combined with the received set of non-default differential gains to generate non-default gains, for example, to compensate for the limitations of the default dynamic range compression curve.

[0023] The techniques described in this document do not require an audio decoder to lock in to audio processing (such as irreversible) that may have been performed by an upstream device such as an audio encoder while assuming a hypothetical playback environment, scenario, etc. in a hypothetical audio decoder. The decoder described in this document may be configured to customize audio processing operations based on a particular playback scenario, for example, to distinguish between various loudness levels present in audio content, minimize loss of audio perceptual quality at or near a boundary loudness level, and maintain spatial balance between channels or a subset of channels.

[0024] An audio decoder that receives an encoded audio signal having a dynamic range compression curve, a reference loudness level, an attack time, a release time, etc. can determine a particular playback environment being used in the decoder and select a particular compression curve having a corresponding reference loudness level corresponding to the particular playback environment.

[0025] The decoder can calculate / determine the loudness level in individual parts of the audio content (such as audio data blocks, audio data frames, etc.) extracted from the encoded audio signal, or obtain the loudness level in individual parts of the audio content if the audio encoder has calculated the loudness level and provided it in the encoded audio signal. Based on one or more of the loudness level in individual parts of the audio content, the loudness level in previous parts of the audio content, the loudness level in subsequent parts of the audio content if available, the specific compression curve, the specific playback environment or scenario-related specific profile, etc., the decoder determines audio processing parameters such as the gain (DRC gain) for dynamic range control, attack time, release time, etc. The audio processing parameters can also include adjustments to align the dialog loudness level to a specific reference loudness level (which may be user-adjustable) for a specific playback environment.

[0026] The decoder applies audio processing operations including dynamic range control (such as multi-channel, multi-band, etc.), dialog level adjustment, etc. with the audio processing parameters. The audio processing operations executed by the decoder may further include, but are not limited to, gain smoothing based on the attack time and release time provided as part of or in relation to a selected dynamic range compression curve, gain limiting to prevent clipping, etc. Different audio processing operations may be executed with different (such as adjustable, threshold-dependent, controllable, etc.) time constants. For example, the gain limiting to prevent clipping may be applied to individual audio data blocks, individual audio data frames, etc. with a relatively short time constant (such as instantaneous, about 5.3 milliseconds, etc.).

[0027] In some embodiments, the decoder extracts ASA parameters (e.g., the temporal position of an auditory event boundary, the time-dependent value of an event certainty index, etc.) from the metadata in the encoded audio signal, and controls the rate of gain smoothing in the auditory event based on the extracted ASA parameters (e.g., using a short time constant for attack at the auditory event boundary, using a long time constant to slow down the gain smoothing within the auditory event, etc.).

[0028] In some embodiments, the decoder also maintains a histogram of the instantaneous loudness levels for a certain time interval or window, and uses the histogram to control the rate of gain change in loudness level transitions (e.g., between programs, between a program and a commercial, etc.) by, for example, modifying the time constant.

[0029] In some embodiments, the decoder supports two or more speaker configurations (e.g., portable mode with speakers, portable mode with headphones, stereo mode, multi-channel mode, etc.). The decoder may be configured to maintain the same loudness level between two different speaker configurations (e.g., between stereo mode and multi-channel mode, etc.) when playing the same audio content, for example. The audio decoder may use one or more downmixing formulas to downmix the multi-channel audio content received from the encoded audio signal for a certain reference speaker configuration to a specific speaker configuration in the audio decoder. The multi-channel audio content is encoded for the reference speaker configuration.

[0030] In some embodiments, automatic gain control (AGC) may be disabled in the audio decoder described in this document.

[0031] In some embodiments, it forms part of a media processing system, including but not limited to, an audio visual device, a flat panel TV, a handheld device, a game console, a television, a home theater system, a tablet, a mobile device, a laptop computer, a netbook computer, a cellular radio phone, an e-book reader, a point of sale terminal, a desktop computer, a computer workstation, a computer kiosk, and various other types of terminals and media processing units.

[0032] Various modifications to the preferred embodiments and general principles and features described herein will be readily apparent to those skilled in the art. Accordingly, the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.

[0033] 〈2. Dynamic Range Control〉 Without customized dynamic range control, input audio information (e.g., PCM samples, time-frequency samples in a QMF matrix, etc.) is often reproduced at a loudness level inappropriate for a particular playback environment of the playback device (i.e., including the physical and / or mechanical playback limitations of the device). This is because the particular playback environment of the playback device may be different from the playback environment targeted when the encoded audio content was encoded at the encoding device.

[0034] The techniques described herein can be used to support customized dynamic range control of a wide variety of audio content for any of a wide variety of playback environments while maintaining the perceptual quality of the audio content.

[0035] Dynamic range control (DRC) refers to a time-dependent audio processing operation that changes the input dynamic range of the loudness level in audio content to an output dynamic range different from the input dynamic range (e.g., compressing, cutting, expanding, boosting, etc.). For example, in a dynamic range control scenario, soft sounds may be mapped (e.g., boosted) to a higher loudness level, and loud sounds may be mapped (e.g., cut) to a lower loudness value. As a result, in the loudness domain, in this example, the output range of the loudness level becomes smaller than the input range of the loudness level. However, in some embodiments, dynamic range control may be reversible such that the original range is restored. For example, as long as the mapped loudness level in the output dynamic range mapped from the original loudness level is below the clipping level and each unique original loudness level is mapped to a unique output loudness level, an expansion operation may be performed to restore the original range.

[0036] The DRC techniques described herein can be used to provide a better listening experience in certain playback environments or situations. For example, soft sounds in a noisy environment may be masked by the noise that makes the soft sounds inaudible. Conversely, in some situations, such as a noisy neighbor, loud sounds may not be desired. Typically, many devices with small form factor loudspeakers cannot reproduce sound at high output levels. In some cases, lower signal levels may be reproduced below the human auditory threshold. The DRC techniques can perform mapping the input loudness level to the output loudness level based on the DRC gain (e.g., a scaling factor for scaling the audio amplitude, a boost ratio, a cut ratio, etc.) found using a dynamic range compression curve.

[0037] A dynamic range compression curve is a function (such as a look-up table, a curve, a multi-segment piecewise linear line, etc.) that maps individual input loudness levels (such as those from individual audio data frames, for example, sounds other than dialog) determined from individual audio data frames to individual gains or gains for dynamic range control. Each of the individual gains indicates the magnitude of the gain to be applied to the corresponding individual input loudness level. The output loudness level after applying the individual gains represents the target loudness level for the audio content in the individual audio data frame in a specific playback environment.

[0038] In addition to specifying the mapping between the gain and the loudness level, the dynamic range compression curve may include or be provided with specific release times and attack times when applying a specific gain. Attack refers to the increase in signal energy (or loudness) between successive time samples. On the other hand, release refers to the decrease in signal energy (or loudness) between successive time samples. The attack time (for example, 10 milliseconds, 20 milliseconds, etc.) refers to the time constant used to smooth the DRC gain when the corresponding signal is in the attack mode. The release time (for example, 80 milliseconds, 100 milliseconds, etc.) refers to the time constant used to smooth the DRC gain when the corresponding signal is in the release mode. In some embodiments, additionally, optionally, or alternatively, these time constants are used for smoothing the signal energy (loudness) before determining the DRC gain.

[0039] Different playback environments may correspond to different dynamic range compression curves. For example, the dynamic range compression curve for the playback environment of a flat panel TV may be different from the dynamic range compression curve for the playback environment of a portable device. In some embodiments, the playback device may have two or more playback environments. For example, the first dynamic range compression curve for the first playback environment of a portable device using speakers may be different from the second dynamic range compression curve for the second playback environment of the same portable device using headphones.

[0040] 〈3. Audio Decoder〉 FIG. 1A shows an exemplary audio decoder 100 having a data extractor 104, a dynamic range controller 106, an audio renderer 108, and the like.

[0041] In some embodiments, the data extractor (104) is configured to receive an encoded input signal 102. The encoded input signal as described herein may be a bitstream that includes an encoded (e.g., compressed, etc.) input audio data frame and metadata. The data extractor (104) is configured to extract / decode the input audio data frame and metadata from the encoded input signal (102). Each of the input audio data frames has a plurality of encoded audio data blocks, each of which represents a plurality of audio samples. Each frame represents a (e.g., fixed) time interval that includes a certain number of audio samples. The frame size can vary with the sample rate and the encoded data rate. The audio samples are quantized audio data elements (e.g., input PCM samples, input time - frequency samples in a QMF matrix, etc.) that represent spectral content in one, two, or more (audio) frequency bands or frequency ranges. The quantized audio data elements in the input audio data frame may represent a pressure wave in the digital (quantized) domain. The quantized audio data elements can cover a finite range of loudness levels below a maximum possible value (e.g., clipping level, maximum loudness level, etc.).

[0042] Metadata can be used by a wide variety of receiving - side decoders for processing input audio - data frames. The metadata may include various operation parameters related to one or more operations to be performed by the decoder (100), normalization parameters related to the dialog loudness level represented in the input audio - data frame, and the like. The dialog loudness level may refer to levels (such as psycho - acoustic, perceptual, etc.) of dialog loudness, program loudness, average dialog loudness, etc. in an entire program (such as a movie, a TV program, a radio broadcast, etc.), a part of the program, the dialog of the program, and the like.

[0043] Some or all of the operations and functions of the decoder (100) or its modules (such as the data extractor 104, the dynamic - range controller 106, etc.) may be adapted in response to metadata extracted from the encoded input signal (102). For example, metadata - including but not limited to dynamic - range compression curves, dialog loudness levels, etc. - may be used by the decoder (100) to generate digital - domain output audio - data elements (such as output PCM samples, output time - frequency samples in a QMF matrix, etc.). The output data elements can then be used to drive audio channels or speakers to achieve a specified loudness or reference playback level during playback in a particular playback environment.

[0044] In some embodiments, the dynamic - range controller (106) is configured to receive some or all of the audio - data elements and metadata in the input audio - data frame and perform audio - processing operations (such as dynamic - range control operations, gain - smoothing operations, gain - limiting operations, etc.) on the audio - data elements in the input audio - data frame based at least in part on metadata extracted from the at least partially encoded audio signal (102).

[0045] In some embodiments, the dynamic range controller (106) may include a selector 110, a loudness calculator 112, a DRC gain unit 114, etc. The selector (110) may be configured to determine a speaker configuration related to a specific playback environment in the decoder (100) (e.g., flat panel mode, portable device with speakers, portable device with headphones, 5.1 speaker configuration, 7.1 speaker configuration, etc.) and select a specific dynamic range compression curve from the dynamic range compression curves extracted from the encoded input signal (102), etc.

[0046] The loudness calculator (112) may be configured to calculate one or more types of loudness levels represented by audio data elements in the input audio data frame. Exemplary types of loudness levels include, but are not limited to, individual loudness levels over individual frequency bands in individual channels over individual time intervals, broadband (or wideband) loudness levels over wide (or broad) frequency ranges in individual channels, loudness levels determined from or smoothed over an audio data block or frame, loudness levels determined from or smoothed over two or more audio data blocks or frames, loudness levels smoothed over one or more time intervals, etc. Zero, one, or more of these loudness levels may be modified by the decoder (100) for dynamic range control.

[0047] To determine the loudness level, the loudness calculator (112) can determine one or more time-dependent physical sound wave attributes such as the spatial pressure level at a particular audio frequency, represented by audio data elements in an input audio data frame. The loudness calculator (112) can use the one or more time-varying physical wave attributes to derive one or more types of loudness levels based on one or more psychoacoustic functions that model human loudness perception. The psychoacoustic function may be a non-linear function that converts a particular spatial pressure level at a particular audio frequency to a particular loudness for that particular audio frequency - constructed based on a model of the human auditory system - and the like.

[0048] The loudness levels over a plurality of (audio frequencies) or over a plurality of frequency bands (such as broadband, wideband, etc.) may be derived through the integration of the specific loudness levels over the plurality of (audio) frequencies or over the plurality of frequency bands. The time-averaged, smoothed, etc. loudness levels over one or more time intervals (such as longer than represented by audio data elements in an audio data block or frame) may be obtained using one or more smoothing filters implemented as part of the audio processing operations in the decoder (100).

[0049] In one exemplary embodiment, specific loudness levels for different frequency bands may be calculated for each audio data block of a certain number (such as 256) of samples. A pre-filter may be used to apply frequency weighting (such as something similar to IEC B weighting) to the specific loudness levels in order to integrate the specific loudness levels into a broadband loudness level. The sum of the wide loudness levels across two or more channels (such as left front, right front, center, left surround, right surround, etc.) may be performed to provide the overall loudness level of the two or more channels.

[0050] In some embodiments, the overall loudness level may refer to the broadband loudness level in a certain channel (such as the center) of a speaker configuration. In some embodiments, the overall loudness level may refer to the broadband (or wideband) loudness levels in a plurality of channels. The plurality of channels may be all channels in a speaker configuration. Additionally, optionally or alternatively, the plurality of channels may include a subset of channels in a speaker configuration (such as a subset of channels including left front, right front and low frequency effects (LFE), a subset of channels including left surround and right surround, a subset of channels including the center, etc.).

[0051] The loudness level (e.g., broadband, wideband, overall, specific, etc.) may be used as an input for finding the corresponding DRC gain (e.g., static, pre-smoothed, pre-limited, etc.) from the selected dynamic range compression curve. The loudness level used as an input for finding the DRC gain may first be adjusted or normalized with respect to the dialog loudness level from the metadata extracted from the encoded audio signal (102). In some embodiments, the adjustment and normalization related to the adjustment of the dialog loudness level may be performed, but not limited to, on a part of the audio content in the encoded audio signal (102) in a non-loudness region (e.g., SPL region, etc.) before a specific spatial pressure level represented in the part of the audio content in the encoded audio signal (102) is converted or mapped to a specific loudness level of the part of the audio content in the encoded audio signal (102).

[0052] In some embodiments, the DRC gain unit (114) is configured with a DRC algorithm to generate a gain (e.g., for dynamic range control, for gain limiting, for gain smoothing, etc.), and apply the gain to one or more loudness levels at one or more types of loudness levels represented by audio data elements in an input audio data frame to achieve a target loudness level for that particular playback environment. The application of a gain (such as a DRC gain as described herein) is not essential but may occur in the loudness domain. In some embodiments, the gain may be generated, smoothed, and applied directly to the input signal based on a loudness calculation (which may be represented in SPL compensated for a zone or simply a dialog loudness level, for example, without conversion). In some embodiments, the techniques as described herein may apply a gain to a signal in the loudness domain, then convert the signal from the loudness domain to the original (linear) SPL domain, and calculate the corresponding gain to be applied to the signal by evaluating the signal before and after the gain is applied to the signal in the loudness domain. Then, the ratio (or difference when represented in logarithmic dB) determines the corresponding gain for that signal.

[0053] In some embodiments, the DRC algorithm operates with a plurality of DRC parameters. The DRC parameters are already calculated by an upstream encoder (such as 150 etc.) and embedded in the encoded audio signal (102), and can be obtained by the decoder (100) from the metadata in the encoded audio signal (102), including the dialog loudness level. The dialog loudness level from the upstream encoder indicates an average dialog loudness level (such as for each program, for the energy of a reference rectangular wave relative to the energy of a full-scale 1 kHz sine wave, etc.). In some embodiments, the dialog loudness level extracted from the encoded audio signal (102) may be used to reduce the difference in loudness levels between programs. In one embodiment, the reference dialog loudness level may be set to the same value between different programs in the same specific playback environment in the decoder (100). Based on the dialog loudness level from the metadata, such that the output dialog loudness level averaged over a plurality of audio data blocks of a program is raised / lowered to the reference dialog loudness level for that program (such as pre-configured, system default, user-configurable, profile-dependent, etc.), the DRC gain unit (114) can apply a dialog loudness-related gain to each audio data block in the program.

[0054] In some embodiments, the DRC gain may be used to address differences in loudness levels within a program by boosting or cutting signal portions in soft and / or loud sounds according to a selected dynamic range compression curve. One or more of these DRC gains may be calculated / determined by a DRC algorithm based on a selected dynamic range compression curve and a loudness level (such as broadband, wideband, overall, specific, etc.) determined from one or more of the corresponding audio data blocks, audio data frames, etc.

[0055] The loudness level used to determine the DRC gain by searching for a selected dynamic range compression curve (such as static, unsmoothed, pre-gain-limited, etc.) may be calculated over a short interval (such as about 5.3 milliseconds, etc.). The integration time of the human auditory system (such as about 200 milliseconds, etc.) can be much longer. The DRC gain obtained from the selected dynamic range compression curve may be smoothed with a certain time constant to account for the long integration time of the human auditory system. To implement a fast rate of change (increase or decrease) in the loudness level, a short time constant may be used to cause a change in the loudness level in a short time interval corresponding to the short time constant. Conversely, to implement a slow rate of change (increase or decrease) in the loudness level, a long time constant may be used to change the loudness level in a long time interval corresponding to the long time constant.

[0056] The human auditory system may respond with different integration times to increasing and decreasing loudness levels. In some embodiments, different time constants may be used depending on whether the loudness level is increasing or decreasing in order to smooth the static DRC gain retrieved from a selected dynamic range compression curve. For example, corresponding to the characteristics of the human auditory system, attack (increase in loudness level) is smoothed with a relatively short time constant (such as attack time), while release (decrease in loudness level) is smoothed with a relatively long time constant (such as release time).

[0057] The DRC gain for a portion of the audio content (such as one or more of an audio data block, an audio data frame, etc.) may be calculated using the loudness level determined from the said portion of the audio content. The loudness level to be used for the search in the selected dynamic range compression curve may first be adjusted with respect to (such as in relation to, etc.) the dialog loudness level in the metadata extracted from the encoded audio signal (102) (such as for a program of which the audio content forms a part).

[0058] The reference dialog loudness level (such as -31 dB in "Line" mode FS , -20 dB in "RF" mode FS etc.) may be specified or established for a particular playback environment in the decoder (100). Additionally, alternatively or optionally, in some embodiments, the user may be given control to set or change the reference dialog loudness level in the decoder (100).

[0059] The DRC gain unit (114) can be configured to determine a dialog loudness related gain for audio content to cause a change from an input dialog loudness level to a reference dialog loudness level as the output dialog loudness level.

[0060] In some embodiments, the DRC gain unit (114) may be configured to handle peak levels in a particular playback environment in the decoder (100) and adjust the DRC gain to prevent clipping. In some embodiments, under a first approach, if the audio content extracted from the encoded audio signal (102) includes audio data elements for a reference multi-channel configuration having more channels than the channels of a particular speaker configuration in the decoder, a particular speaker configuration downmix from the reference multi-channel configuration may be performed before discriminating and processing the peak level for clipping prevention. Additionally, optionally or alternatively, in some embodiments, under a second approach, if the audio content extracted from the encoded audio signal (102) includes audio data elements for a reference multi-channel configuration having more channels than the channels of a particular speaker configuration in the decoder, a downmix formula (e.g., ITU stereo downmix, matrixed-surround compatible downmix, etc.) may be used to obtain the peak level for a particular speaker configuration in the decoder (100). The peak level may be adjusted to reflect the change from the input dialog loudness level to a reference dialog loudness level as the output dialog loudness level. The maximum allowable gain that does not cause clipping (for a certain audio data block, for a certain audio data frame, etc.) may be determined based at least in part on the reciprocal of the peak level (e.g., multiplied by -1, etc.). Thus, the audio decoder under the techniques described herein can be configured to accurately determine the peak level and apply clipping prevention specifically for the decoder-side playback configuration. Neither the audio decoder nor the audio encoder needs to make assumptions about a worst-case scenario in a hypothetical decoder.In particular, the decoder in the above first approach can accurately determine the peak level and apply post-downmix clipping prevention without using the downmix formula, downmix channel gain, etc. (which are used under the second approach as described above).

[0061] In some implementations, a combination of adjustments to the dialog loudness level and DRC gain can prevent peak level clipping even in the worst-case downmix (e.g., one that produces the maximum peak level after downmixing, one that produces the maximum downmix channel gain, etc.). However, in some other embodiments, a combination of adjustments to the dialog loudness level and DRC gain may not be sufficient to prevent peak level clipping. In these embodiments, the DRC gain may be replaced (e.g., capped, etc.) by the highest gain that prevents clipping at the peak level.

[0062] In some embodiments, the DRC gain unit (114) is configured to obtain a time constant (e.g., attack time, release time, etc.) from metadata extracted from the encoded audio signal (102). The DRC gain, time constant, maximum allowable gain, etc. may be used by the DRC gain unit (114) to perform DRC, gain smoothing, gain limiting, etc.

[0063] For example, the application of the DRC gain may be smoothed with a filter controlled by a certain time constant. The gain limiting operation may be implemented by a min() function that takes the smaller of the gain to be applied and the maximum allowable gain for that gain. Through this function, the gain (e.g., before limiting, such as DRC) may be immediately replaced by the maximum allowable gain over a relatively short time interval, etc., thereby preventing clipping.

[0064] In some embodiments, the audio renderer (108) applies a gain determined based on, for example, DRC, gain limiting, gain smoothing, etc. to the input audio data extracted from the encoded audio signal (102), and then generates channel-specific audio data (116) (such as for a multi-channel, etc.) for that particular speaker configuration. The channel-specific audio data (118) may be used to drive the speakers, headphones, etc. represented in the speaker configuration.

[0065] Additionally and / or optionally, in some embodiments, the decoder (100) may be configured to perform one or more other operations related to pre-processing, post-processing, rendering, etc. related to the input audio data.

[0066] The techniques described herein can be used with a variety of different speaker configurations (such as 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10 - 60 speaker configurations, 60+ speaker configurations, object signals or combinations of object signals, etc.) corresponding to a variety of different surround sound configurations and a variety of different rendering environment configurations (such as movie theaters, parks, opera houses, concert halls, bars, homes, lecture halls, etc.).

[0067] 〈4. Audio Encoder〉 Figure 1B shows an exemplary encoder 150. The encoder (150) may have an audio content interface 152, a dialogue loudness analyzer 154, a DRC reference storage 156, an audio signal encoder 158, etc. The encoder 150 may be part of a broadcast system, an Internet-based content server, an over-the-air network operator system, a movie production system, etc.

[0068] In some embodiments, the audio content interface (152) is configured to receive, for example, audio content 160, audio content control inputs 162, etc., and to generate an encoded audio signal (e.g., 102) based at least in part on the audio content (160), a part or all of the audio content control inputs (162). For example, the audio content interface (152) may be used to receive the audio content (160), audio content control inputs (162) from a content creator, content provider, etc.

[0069] The audio content may form part or all of the overall media data including only audio, audio-visual, etc. The audio content (160) may include one or more of portions of a program, a program, several programs, one or more commercials, etc.

[0070] In some embodiments, the dialog loudness analyzer (154) is configured to determine / establish one or more dialog loudness levels of one or more parts of the audio content (152) (e.g., one or more programs, one or more commercials, etc.). In some embodiments, the audio content is represented by one or more sets of audio tracks. In some embodiments, the dialog audio content of the audio content is on a separate audio track. In some embodiments, at least a part of the audio content is on an audio track including non-dialog audio content.

[0071] The audio content control input (162) may include some or all of a user control input, a control input provided by an external system / device for the encoder (150), a control input from a content creator, a control input from a content provider, etc. For example, a user such as a mixing engineer can provide / specify one or more dynamic range compression curve identifiers. Those identifiers may be used to retrieve one or more dynamic range compression curves that best fit the audio content (160) from a data storage unit such as the DRC reference storage unit (156).

[0072] In some embodiments, the DRC reference storage unit (156) is configured to store a set of DRC reference parameters, etc. Those sets of DRC reference parameters may include definition data about one or more dynamic range compression curves, etc. In some embodiments, the encoder (150) may encode two or more dynamic range compression curves (e.g., simultaneously in parallel, etc.) into the encoded audio signal (102). Zero, one, or more of those dynamic range compression curves may be standard-based, unique, customized, modifiable by a decoder, etc. In an exemplary embodiment, both of the dynamic range compression curves of FIGS. 2A and 2B can be encoded (e.g., simultaneously in parallel, etc.) into the encoded audio signal (102).

[0073] In some embodiments, the audio signal encoder (158) receives audio content from the audio content interface (152), a dialog loudness level from the dialog loudness analyzer (154), etc., retrieves one or more sets of DRC reference parameters from the DRC reference storage unit (156), formats the audio content into audio data blocks / frames, formats the dialog loudness level, the set of DRC reference parameters, etc. into metadata (e.g., metadata container, metadata field, metadata structure, etc.), and can be configured to encode the audio data blocks / frames and the metadata into an encoded audio signal (102).

[0074] The audio content to be encoded in the encoded audio signal as described herein can be received in one or more of a variety of ways, such as wirelessly, via a wired connection, through a file, via an Internet download, etc., in one or more of a variety of source audio formats.

[0075] The encoded audio signal described herein can be part of an overall media data bitstream (e.g., for an audio broadcast, an audio program, an audiovisual program, an audiovisual broadcast, etc.). The media data bitstream can be accessed from a server, a computer, a media storage device, a media database, a media file, etc. The media data bitstream may be broadcast, transmitted, or received through one or more wireless or wired network links. The media data bitstream may be communicated through one or more such media as a network connection, a USB connection, a wide area network, a local area network, a wireless connection, an optical connection, a bus, a crossbar connection, a serial connection, etc.

[0076] Any of the components depicted (e.g., in FIGS. 1A, 1B, etc.) may be implemented as one or more processes and / or one or more IC circuits (e.g., ASIC, FPGA, etc.) in hardware, software, or a combination of hardware and software.

[0077] 〈5. Dynamic Range Compression Curve〉 FIGS. 2A and 2B show exemplary dynamic range compression curves that can be used by a DRC gain unit (104) in a decoder (100) to derive a DRC gain from an input loudness level. As shown in the figures, the dynamic range compression curve may be centered on a reference loudness level in a program to provide an overall gain appropriate for a particular playback environment. Exemplary definition data for the dynamic range compression curve (e.g., within the metadata of the encoded audio signal 102) (e.g., including, but not limited to, boost ratio, cut ratio, attack time, release time, etc.) is shown in the table below. Here, each profile in a plurality of profiles (e.g., film standard, film light, music standard, music light, speech, etc.) represents a particular playback environment (e.g., in a decoder 100, etc.).

[0078] [Table 1] Some embodiments may receive one or more compression curves described using loudness levels represented in dB SPL or dB FS and gains represented in dB with respect to dB SPL On the other hand, the DRC gain is dB SPLIt is executed with different loudness representations (e.g., zones) that have a non-linear relationship with the loudness level. At this time, the compression curve used in the DRC gain calculation may be converted as described using the different loudness representations (e.g., zones).

[0079] 〈6. DRC Gain, Gain Limitation and Gain Smoothing〉 Figure 3 shows exemplary processing logic for determining / calculating combined DRC and limiting gains. The processing logic may be implemented by a decoder (100), an encoder (150), etc. For illustration purposes only, a DRC gain unit (e.g., 114) in a decoder (e.g., 100, etc.) may be used to implement the processing logic.

[0080] The DRC gain for a portion of the audio content (e.g., one or more of an audio data block, an audio data frame, etc.) may be calculated using the loudness level determined from that portion of the audio content. The loudness level may first be adjusted with respect to (e.g., in relation to, etc.) the dialog loudness level in the metadata extracted from the encoded audio signal (102) (e.g., for a program of which the audio content is a part). In the example shown in Figure 3, the difference between the loudness level of the portion of the audio content and the dialog loudness level ("dialnorm") may be used as an input for finding the DRC gain from the selected dynamic range compression curve.

[0081] To prevent clipping of the output audio data elements in that particular playback environment, the DRC gain unit (114) may be configured to handle the peak levels in a particular playback scenario (e.g., specific to a particular combination of the encoded audio signal 102 and the playback environment in the decoder 100, etc.). The playback scenario may be one of a variety of possible playback scenarios (e.g., a multi-channel scenario, a downmix scenario, etc.).

[0082] In some embodiments, the individual peak levels for individual portions of audio content (e.g., audio data blocks, several audio data blocks, audio data frames, etc.) at a particular time resolution may be provided as part of the metadata extracted from the encoded audio signal (102).

[0083] In some embodiments, the DRC gain unit (114) can be configured to determine peak levels in these scenarios and adjust the DRC gain if necessary. During the calculation of the DRC gain, a parallel process may be used by the DRC gain unit (114) to determine the peak level of the audio content. For example, the audio content may be encoded for a reference multi-channel configuration having more channels than the channels of a particular speaker configuration used by the decoder (100). The audio content for the additional channels of the reference multi-channel configuration may be converted to downmixed audio data (e.g., ITU stereo downmix, matrixed-surround compatible downmix, etc.) to derive fewer channels for a particular speaker configuration in the decoder (100). In some embodiments, under a first approach, the downmix from the reference multi-channel configuration to a particular speaker configuration may be performed before determining and processing the peak level for clipping prevention. Additionally, optionally, or alternatively, in some embodiments, under a second approach, the downmix channel gain associated with downmixing the audio content may be used as part of the input for adjusting, deriving, calculating, etc., the peak level for that particular speaker configuration. In an exemplary embodiment, the downmix channel gain may be derived at least in part based on one or more downmix equations used to perform the downmix operation from the reference multi-channel configuration to a particular speaker configuration in the playback environment in the decoder (100).

[0084] In some media applications, the reference dialog loudness level (e.g., -31 dB in "Line" mode FS , -20 dB in "RF" mode FSFor example, etc.) may be specified or assumed for a specific playback environment in the decoder (100). In some embodiments, the user may be given control over setting or changing the reference dialog loudness level in the decoder (100).

[0085] A dialog loudness gain may be applied to the audio content to adjust the (e.g., output) dialog loudness level to the reference dialog loudness level. To reflect this adjustment, the peak level should be adjusted appropriately. In one example, the (input) dialog loudness level is -23dB FS may be. If the reference dialog loudness level is -31dB FS In the "line" mode of, to produce the output dialog loudness level of the reference dialog loudness level, the adjustment to the (input) dialog loudness level is -8dB. In this "line" mode, the adjustment to the peak level is also -8dB, which is the same as the adjustment to the dialog loudness level. If the reference dialog loudness level is -20dB FS In the "RF" mode of, to produce the output dialog loudness level of the reference dialog loudness level, the adjustment to the (input) dialog loudness level is 3dB. In this "RF" mode, the adjustment to the peak level is also 3dB, which is the same as the adjustment to the dialog loudness level.

[0086] The sum of the difference between the peak level and the reference dialog loudness level (denoted as "dialref") and the dialog loudness level ("dialnorm") in the metadata from the encoded audio signal (102) may be used as an input for calculating the maximum (e.g., allowed, etc.) gain for the DRC gain. The adjusted peak level is (0dB FS with respect to the clipping level of) FSSince it is represented by, for example, the maximum allowable gain that does not cause clipping (for example, for the current audio data block, for the current audio data frame, etc.) is simply the reciprocal of the adjusted peak level (for example, multiplied by -1, etc.).

[0087] In some embodiments, even if the dynamic range compression curve from which the DRC gain is derived is designed to cut loud sounds to some extent, the peak level may exceed the clipping level (represented by 0 dB). In some embodiments, the combination of the dialog loudness level and the adjustment to the DRC gain prevents peak level clipping even in the worst-case downmix (for example, the one that generates the maximum downmix channel gain, etc.). However, in some other embodiments, the combination of the dialog loudness level and the adjustment to the DRC gain may not be sufficient to prevent peak level clipping. In these embodiments, the DRC gain may be replaced (for example, capped, etc.) by the highest gain that prevents clipping at the peak level. FS In some embodiments, the DRC gain unit (114) is configured to obtain time constants (for example, attack time, release time, etc.) from metadata extracted from the encoded audio signal (102). These time constants may or may not change along with one or more of the dialog loudness level of the audio content or the current loudness level. The DRC gain, time constants, and maximum gain retrieved from the dynamic range compression curve may be used to perform gain smoothing and limiting operations.

[0088] In some embodiments, the DRC gain unit (114) is configured to obtain time constants (for example, attack time, release time, etc.) from metadata extracted from the encoded audio signal (102). These time constants may or may not change along with one or more of the dialog loudness level of the audio content or the current loudness level. The DRC gain, time constants, and maximum gain retrieved from the dynamic range compression curve may be used to perform gain smoothing and limiting operations.

[0089] In some embodiments, potentially gain-limited DRC gain does not exceed the maximum peak loudness level in a particular playback environment. The static DRC gain derived from the loudness level may be smoothed with a filter controlled by a time constant. The limiting operation may be implemented by one or more min() functions. Through this function, the DRC gain (e.g., before limiting) may be immediately replaced by the maximum allowable gain over a relatively short time interval, etc., thereby preventing clipping. The DRC algorithm may be configured to smoothly release from the clipping gain to a lower gain as the peak level of the incoming audio content transitions from above the clipping level to below the clipping level.

[0090] To perform the determination / calculation / application of the DRC gain shown in FIG. 3, one or more different implementations (e.g., real-time, two-pass, etc.) may be used. By way of example only, the adjustment to the dialog loudness level, the DRC gain (e.g., static, etc.), the time-dependent gain variation due to smoothing, the gain clipping due to limiting, etc., has been described as a combined gain from the above DRC algorithm. However, in various embodiments, for controlling the dialog loudness level (e.g., between different programs, etc.), for dynamic range control (e.g., for different parts of the same program, etc.), for preventing clipping, for gain smoothing, etc., other approaches for applying gain to the audio content may be used. For example, some or all of the adjustment to the dialog loudness level, the DRC gain (e.g., static, etc.), the time-dependent gain variation due to smoothing, the gain clipping due to limiting, etc., can be applied partially / individually, applied serially, applied in parallel, applied partially serially and partially in parallel, etc.

[0091] 〈7. Input Smoothing and Gain Smoothing〉 In addition to DRC gain smoothing, in various embodiments, other smoothing processes may be implemented under the techniques described herein. In one example, input smoothing may be used, and input audio data extracted from the encoded audio signal (102) may be smoothed using, for example, a simple single-pole smoothing filter to obtain a spectrum of a particular loudness level with better temporal characteristics (e.g., more temporally smooth, fewer temporal spikes, etc.) than the spectrum of the particular loudness level without input smoothing.

[0092] In some embodiments, the different smoothing processes described herein can use different time constants (e.g., 1 second, 4 seconds, etc.). In some embodiments, two or more smoothing processes can use the same time constant. In some embodiments, the time constant used in the smoothing processes described herein may be frequency-dependent. In some embodiments, the time constant used in the smoothing processes described herein may be frequency-independent.

[0093] One or more smoothing processes may be connected to a reset process that supports automatic or manual reset of the one or more smoothing processes. In some embodiments, when a reset occurs in the reset process, the smoothing process may speed up the smoothing operation by switching or transitioning to a smaller time constant. In some embodiments, when a reset occurs in the reset process, the memory of the smoothing process may be reset to a value. This value may be the last input sample to the smoothing process.

[0094] 〈8. DRC across Multiple Frequency Bands〉 In some embodiments, specific loudness levels in specific frequency bands can be used to derive corresponding DRC gains in those specific frequency bands. However, this can lead to a change in the tone color. Those specific loudness levels can vary significantly in different bands, so even when the broadband (or wideband) loudness level across the entire frequency band remains constant, different DRC gains may be applied because of this.

[0095] In some embodiments, instead of applying a DRC gain that varies with each individual frequency band, a DRC gain that does not vary with the frequency band but varies with time is applied instead. The same time-varying DRC gain is applied across all frequency bands. The time-averaged DRC gain of the time-varying DRC gain may be set to be the same as the static DRC gain derived from the selected dynamic range compression curve based on the broadband (or wideband) range or the broadband, wideband, and / or overall loudness level across a plurality of frequency bands. As a result, it is possible to prevent changes to the tone color effects that can be caused by applying different DRC gains in different frequency bands in other approaches.

[0096] In some embodiments, the DRC gain in an individual frequency band is controlled using a broadband (or wideband) DRC gain determined based on a broadband (or wideband) loudness level. The DRC gain in an individual frequency band may operate around the broadband (or wideband) DRC found in the dynamic range compression curve based on the broadband (or wideband) loudness level. Thus, the DRC gain in an individual frequency band time-averaged over a time interval (e.g., longer than 5.3 milliseconds, 20 milliseconds, 50 milliseconds, 80 milliseconds, 100 milliseconds, etc.) is the same as the broadband (wideband) level shown in the dynamic range compression curve. In some embodiments, loudness level fluctuations over a short time interval relative to the time interval, deviating from the time-averaged DRC gain, are acceptable among channels and / or frequency bands. This approach ensures the application of the correct multi-channel and / or multi-band time-averaged DRC gain shown in the dynamic range compression curve and prevents the DRC gain in a short time interval from deviating too much from such time-averaged DRC gain shown in the dynamic range compression curve.

[0097] 〈9. Volume Adjustment in Loudness Region〉 Applying linear processing for volume adjustment to an audio excitation signal under other approaches that do not implement the techniques described in this document may make low audible signal levels inaudible (e.g., fall below the frequency-dependent auditory threshold of the human auditory system).

[0098] Under the techniques described in this document, volume adjustment of audio content is in the physical domain (e.g., dB SPLRather than being in a physical domain (e.g., having a physical representation), it can be made or implemented in a loudness domain (e.g., having a zone representation). In some embodiments, to maintain the perceptual quality and / or integrity of the loudness-level relationships across all bands at all volume levels, the loudness levels of all bands are scaled with the same factor in the loudness domain. The volume adjustment based on setting and adjusting gains in the loudness domain, as described herein, is transformed back into a non-linear process in a physical domain (or in a digital domain representing the physical domain) that applies different scaling factors to audio excitation signals in different frequency bands, and may be implemented through such non-linear processing. The non-linear processing in the physical domain transformed from the volume adjustment in the loudness domain under the techniques described herein attenuates or improves the loudness level of audio content with a DRC gain that prevents most or all of the low audible levels in the audio content from becoming inaudible. In some embodiments, the difference in loudness levels between loud and soft sounds within a program is reduced - but not perceptually lost - using these DRC gains that keep low audible signal levels above the auditory threshold of the human auditory system. In some embodiments, to maintain similarities such as spectral perception and perceived timbre over a large range of volume levels, at low volume levels, frequencies or frequency bands having excitation signal levels close to the auditory threshold are attenuated less, and thus are perceptually audible.

[0099] The techniques described herein may implement conversions (e.g., to-and-fro conversions, etc.) between signal levels, gains, etc. in a physical domain (or a digital domain representing the physical domain) and loudness levels, gains, etc. in a loudness domain. These conversions may be based on forward and inverse transformation versions of one or more non-linear functions (e.g., mappings, curves, piecewise linear segments, look-up tables, etc.) constructed based on a model of the human auditory system.

[0100] 〈10. Gain Profile by Differential Gain〉 In some embodiments, the audio encoder (e.g., 150, etc.) described in this document is configured to provide profile-related metadata to a downstream audio decoder. For example, the profile-related metadata may be carried in the encoded audio signal as part of the audio-related metadata together with the audio content.

[0101] The profile-related metadata described in this document includes, but is not limited to, definition data for a plurality of gain profiles. One or more first gain profiles (referred to as one or more default gain profiles) in the plurality of gain profiles are represented by one or more corresponding DRC curves (referred to as one or more default DRC curves). The definition data is included in the profile-related metadata. One or more second gain profiles (referred to as one or more non-default gain profiles) in the plurality of gain profiles are represented by one or more corresponding sets of differential gains with respect to the one or more default DRC curves. The definition data is included in the profile-related metadata. More specifically, the default DRC curve (e.g., in the profile-related metadata, etc.) can be used to represent the default gain profile, and the set of differential gains with respect to the default gain profile (e.g., in the profile-related metadata, etc.) can be used to represent the non-default gain profile.

[0102] In some embodiments, the set of differential gains representing a non-default gain profile in relation to a default DRC curve representing a default gain profile includes a gain difference (or gain adjustment) between a set of non-differential (e.g., non-default, etc.) gains generated for the non-default gain profile and a set of non-differential (e.g., default, etc.) gains generated for the default gain profile. Examples of non-differential gains include, but are not limited to, null gain, DRC gain or attenuation, gain or attenuation related to dialog normalization, gain or attenuation related to gain limiting, gain or attenuation related to gain smoothing, etc. The gains described in this document (e.g., non-differential gains, differential gains, etc.) may be time-dependent and may have values that change over time.

[0103] To generate a set of non-differential gains for a gain profile (e.g., default gain profile, non-default gain profile, etc.), the audio encoder described in this document may perform a set of gain generation operations specific to the gain profile. The set of gain generation operations may include DRC operations, gain limiting operations, gain smoothing operations, etc. This may include any of the operations, but is not limited to, operations that are (1) globally applicable to all gain profiles; (2) specific to one or more but not all gain profiles, specific to one or more default DRC curves; (3) specific to one or more non-default DRC curves; (4) specific to the corresponding (e.g., default, non-default, etc.) gain profile; (5) related to one or more algorithms, curves, functions, operations, parameters, etc. that exceed the limits of parameterization supported by media encoding formats, media standards, media-specific specifications, etc.; (6) related to one or more algorithms, curves, functions, operations, parameters, etc. that are not yet generally implemented in available audio decoding devices.

[0104] In some embodiments, the audio decoder (150) determines a set of differential gains for audio content (152) based on a default gain profile represented by a default DRC curve (e.g., such as definition data in the profile-related metadata of the encoded audio signal) and a non-default gain profile different from the default gain profile, and configures the set of differential gains to be included as part of the profile-related metadata in the encoded audio signal as a representation of the non-default gain profile (e.g., with respect to the default DRC curve). The set of differential gains extracted from the profile-related metadata in the encoded audio signal in relation to the default DRC curve can be used by the receiving audio decoder to efficiently and consistently perform gain operations (or attenuation operations) in a playback environment or scenario for a specific gain profile represented by the set of differential gains in relation to the default DRC curve. This enables the receiving audio decoder to apply gain or attenuation for that specific gain profile without requiring the receiving audio decoder to implement a set of gain generation operations. To generate the gain or attenuation, the set of gain generation operations can be implemented in the audio encoder (150).

[0105] In some embodiments, one or more sets of differential gains may be included in the profile-related metadata by the audio encoder (150). Each of the one or more sets of differential gains may be derived from corresponding non-default gain profiles in one or more non-default gain profiles in relation to corresponding default gain profiles in one of the one or more default gain profiles. For example, the first set of differential gains in the one or more sets of differential gains may be derived from a first non-default gain profile in relation to a first default gain profile, while the second set of differential gains in those sets of differential gains may be derived from a second non-default gain profile in relation to a second default gain profile.

[0106] In some embodiments, the first set of differential gains includes a first gain difference (or gain adjustment) determined between a first set of non-differential non-default gains generated based on the first non-default gain profile and a first set of non-differential default gains generated based on the first default gain profile. On the other hand, the second set of differential gains includes a second gain difference determined between a second set of non-differential non-default gains generated based on the second non-default gain profile and a second set of non-differential default gains generated based on the second default gain profile.

[0107] The first default gain profile and the second default gain profile may be the same (e.g., represented by the same default DRC curve with the same set of gain generation operations), or they may be different (e.g., represented by different default DRC curves, with different sets of gain generation operations, represented by a certain default DRC, etc.). In various embodiments, additionally, optionally, or alternatively, the first non-default gain profile may or may not be the same as the second non-default gain profile.

[0108] The profile-related metadata generated by the audio encoder (150) can carry one or more specific flags, indicators, data fields, etc. to indicate the existence of one or more sets of differential gains for one or more corresponding non-default gain profiles. The profile-related data may also include preference flags, indicators, data fields, etc. to indicate which non-default gain profile is preferred for rendering the audio content in a particular playback environment or scenario.

[0109] In some embodiments, the audio decoder (such as 100 etc.) described in this document is configured to decode audio content (such as multi-channel etc.) from the encoded audio signal (102), and extract a dialog loudness level (such as "dialnorm" etc.) from the loudness metadata delivered together with the audio content.

[0110] In some embodiments, an audio decoder (e.g., 100, etc.) is configured to perform at least one set of gain generation operations for gain profiles such as the first default profile, the second default profile, etc. For example, an audio decoder (100) decodes an encoded audio signal (102) having a dialog loudness level (e.g., "dialnorm", etc.); performs a set of gain generation operations to obtain a set of non-differential default gains (or attenuations) for a default gain profile represented by a default DRC curve from which defined data can be extracted by the audio decoder (100) from the encoded audio signal (102); applies the set of non-differential default gains (e.g., the difference between a reference loudness level and "dialnorm", etc.) for the default gain profile during decoding to align / adjust the output dialog loudness level of the sound output to the reference loudness level; and so on.

[0111] Additionally, optionally, or alternatively, in some embodiments, an audio decoder (100) is configured to extract at least one set of differential gains from an encoded audio signal (102). The set of differential gains represents a non-default gain profile in relation to a default DRC curve as discussed above as part of the metadata delivered with the audio content. In some embodiments, the profile relationship metadata includes one or more different sets of differential gains, and each of the one or more different sets of differential gains represents a non-default gain profile in relation to a respective default DRC curve representing a default gain profile. The presence of a DRC curve or a set of differential gains in the profile relationship metadata may be indicated by one or more flags, indicators, data fields carried in the profile relationship metadata.

[0112] In response to determining the existence of the one or more sets of differential gains, the audio decoder (100) can determine / select a set of differential gains corresponding to a particular non-default gain profile from among the one or more different sets of differential gains. The audio decoder (100) can further be configured to identify a default DRC curve with respect to which the set of differential gains represents the particular gain profile, from among definition data for one or more different default DRC curves, for example, in profile-related metadata.

[0113] In some embodiments, the audio decoder (100) is configured to perform a set of gain generation operations to obtain a set of non-differential default gains (or attenuations) for the default gain profile. The set of gain generation operations performed by the audio decoder (100) to obtain the set of non-differential default gains based on the default DRC curve may include one or more operations related to one or more of a standard, proprietary specification, etc. In some embodiments, the audio decoder (100) generates a set of non-differential non-default gains for the particular non-default gain profile based on the set of differential gains from which definition data is extracted from profile-related metadata and the set of non-differential default gains generated by the set of gain generation operations based on the default DRC curve; applies the set of non-differential non-default gains (e.g., the difference between a reference loudness level and "dialnorm") for the default gain profile during decoding to align / adjust the output dialog loudness level of the sound output to the reference loudness level; and so on.

[0114] In some embodiments, an audio decoder (100) can perform gain-related operations for one or more gain profiles. The audio decoder (100) can be configured to determine and perform gain-related operations for a particular gain profile based on one or more factors. These factors can include, but are not limited to: user input specifying a preference for a particular user-selected gain profile, user input specifying a preference for a system-selected gain profile, the capabilities of a particular speaker or audio channel configuration used by the audio decoder (100), the capabilities of the audio decoder (100), the availability of profile-related metadata for the particular gain profile, one or more encoder-generated preference flags for the gain profile, etc. In some embodiments, if there are conflicts among these factors, the audio decoder (100) may implement one or more procedural rules or seek further user input to determine or select a particular gain profile.

[0115] 〈11. Additional Operations Related to Gain〉 Under the techniques described herein, other processing such as dynamic equalization, noise compensation, etc. can be performed in the loudness (e.g., perceptual) domain rather than in the physical domain (or the digital domain representing the physical domain).

[0116] In some embodiments, gains from some or all of various processes such as DRC, equalization noise compensation, clipping prevention, gain smoothing, etc. may be combined into the same gain in the loudness region and / or may be applied in parallel. In some other embodiments, gains from some or all of various processes such as DRC, equalization noise compensation, clipping prevention, gain smoothing, etc. may be separate gains in the loudness region and / or may be applied at least partially in series. In some other embodiments, gains from some or all of various processes such as DRC, equalization noise compensation, clipping prevention, gain smoothing, etc. may be applied in sequence.

[0117] 〈12. Specific and Broadband (or Wideband) Loudness Levels〉 One or more audio processing elements, units, components such as transmission filters, auditory filter banks, synthesis filter banks, short-time Fourier transforms, etc. may be used by an encoder or decoder to perform the audio processing operations described herein.

[0118] In some embodiments, one or more transfer filters that model the filtering of the outer and middle ears of the human auditory system may be used to filter incoming audio signals (e.g., encoded audio signal 102, audio content from a content provider, etc.). In some embodiments, an auditory filter bank may be used to model the frequency selectivity and frequency spread of the human auditory system. The excitation signal levels from some or all of these filters may be determined / calculated and smoothed with a frequency-dependent time constant that becomes shorter at higher frequencies in order to model the integration of energy in the human auditory system. Thereafter, a non-linear function (e.g., relationship, curve, etc.) between the excitation signal and a specific loudness level may be used to obtain a profile of the frequency-dependent specific loudness level. The broadband (or wideband) loudness level can be obtained by integrating the specific loudness over various frequency bands.

[0119] A straightforward summation / integration of a specific loudness level (e.g., using equal weights across all frequency bands) may work well for broadband signals. However, such an approach may underestimate the loudness level (e.g., perceptual) for narrowband signals. In some embodiments, specific loudness levels at different frequencies or in different frequency bands are given different weights.

[0120] In some embodiments, the auditory filter bank and / or transfer filter as described above may be replaced by one or more short-time Fourier transforms (STFTs). The responses of the transfer filter and the auditory filter bank may be applied in the fast Fourier transform (FFT) domain. In some embodiments, for example, when one or more (e.g., forward) transfer filters are used in or before the conversion from the physical domain (or digital domain representing the physical domain) to the loudness domain, one or more inverse transfer filters are used. In some embodiments, for example, when STFTs are used instead of the auditory filter bank and / or transfer filter, no inverse transfer filter is used. In some embodiments, the auditory filter bank is omitted; instead, one or more quadrature mirror filters (QMFs) are used. In these embodiments, the spreading effect of the basilar membrane in the model of the human auditory system may be omitted without significantly affecting the matters of the audio processing operations described in this document.

[0121] Under the techniques described in this document, different numbers of frequency bands (e.g., 20 frequency bands, 40 frequency bands, etc.) may be used in different embodiments. Additionally, optionally, or alternatively, different bandwidths may be used in different embodiments.

[0122] 〈13. Individual Gains for Individual Subsets of Channels〉 In some embodiments, when a particular speaker configuration is a multi-channel configuration, an overall loudness level may be obtained by first adding the excitation signals of all channels before conversion from a physical region (or a digital region representing the physical region) to a loudness region. However, applying the same gain to all channels in a particular speaker configuration may not preserve the spatial balance (such as the balance regarding the relative loudness levels between different channels) between different channels of that particular speaker configuration.

[0123] In some embodiments, in order to preserve the spatial balance so that the relative perceived loudness levels between different channels can be optimally or correctly maintained, the respective loudness levels and the corresponding gains obtained based on the respective loudness levels may be determined or calculated for each channel. In some embodiments, the corresponding gains obtained based on the respective loudness levels are not equal to the same overall gain. For example, each of some or all of the corresponding gains may be equal to the overall gain plus a small correction (such as channel-specific).

[0124] In some embodiments, to preserve spatial balance, each loudness level and the corresponding gain obtained based on each loudness level may be determined or calculated for each subset of channels. In some embodiments, the corresponding gain obtained based on each loudness level is not equal to the same overall gain. For example, each of some or all of the corresponding gains may be equal to the overall gain plus a small correction (e.g., channel-specific). In some embodiments, a subset of channels may include two or more channels that form a proper subset of all channels in that particular speaker configuration (e.g., a subset of channels including left front, right front, and low frequency effect (LFE), a subset of channels including left surround and right surround, etc.). The audio content for a subset of channels may form a submix of the overall mix carried in the encoded audio signal (102). The same gain can be applied to the channels within the submix.

[0125] In some embodiments, to relate the signal level in the digital domain to the corresponding physical (e.g., dB SPL such as spatial pressure by) level in the physical domain that is represented by the digital domain in order to generate the actual loudness from a particular speaker configuration (e.g., as actually perceived), one or more calibration parameters may be used. The one or more calibration parameters may be given values specific to the physical sound equipment in a particular speaker configuration.

[0126] 〈14. Auditory Scene Analysis〉 In some embodiments, the encoder described in this document may implement computer-based auditory scene analysis (ASA) to detect auditory event boundaries in audio content (such as those encoded in the encoded audio signal 102), generate one or more ASA parameters, and format the one or more ASA parameters as part of the encoded audio signal (such as 102) to be delivered to a downstream device (such as decoder 100). The ASA parameters may include, but are not limited to, the position of the auditory event boundary, the value of an auditory event certainty metric (further described below), and the like.

[0127] In some implementations, the (e.g., temporal) position of the auditory event boundary may be indicated in the metadata encoded within the encoded audio signal (102). Additionally, optionally, or alternatively, the (e.g., temporal) position of the auditory event boundary may be indicated in the audio data block and / or frame in which the position of the auditory event boundary is detected (e.g., using a flag, data field, etc.).

[0128] As used in this document, an auditory event boundary refers to the point at which a preceding auditory event ends and / or a subsequent auditory event begins. Each auditory event occurs between two successive auditory event boundaries.

[0129] In some embodiments, the encoder (150) is configured to detect an auditory event boundary based on the difference in a particular loudness spectrum between two consecutive (e.g., temporal) audio data frames. Each particular loudness spectrum may include the non-smoothed loudness spectrum calculated from the corresponding audio data frame of those consecutive audio data frames.

[0130] In some embodiments, the specific loudness spectrum N[b,t] may be normalized to obtain the normalized specific loudness spectrum N NORM [b,t] as shown in the following equation.

[0131] N NORM [b,t]=N[b,t] / max b {N[b,t]} (1) where b represents the frequency band, t represents the time or the audio frame index, and max b {N[b,t]} is the maximum specific loudness level across all frequency bands.

[0132] The normalized specific loudness spectra are subtracted from each other and used to derive the difference absolute value sum D[t] as shown in the following equation.

[0133] D[t]=Σ b |N NORM [b,t]-N NORM [b,t-1]| (2) The difference absolute value sum is mapped to an auditory event certainty index having a value range from 0 to 1 as follows.

[0134]

Equation

[0135] In some embodiments, the encoder (150) is configured to detect an auditory event boundary when D[t] (e.g., at a specific t) exceeds D min (e.g., at the specific t).

[0136] In some embodiments, a decoder (e.g., 100) described herein extracts ASA parameters from an encoded audio signal (e.g., 102) and uses the ASA parameters to prevent unintentional boosting of soft sounds and / or unintentional cutting of loud sounds that cause perceptual distortion of auditory events.

[0137] The decoder (100) may be configured to reduce or prevent unintentional distortion of auditory events by ensuring that the gain is closer to being constant within an auditory event and constraining much of the gain change near the auditory event boundary. For example, the decoder (100) may be configured to use a relatively small time constant (e.g., one comparable to or shorter than the minimum duration of auditory events) in response to a gain change at an attack (e.g., an increase in loudness level) at the auditory event boundary. Thus, the gain change at the attack can be implemented relatively quickly by the decoder (100). On the other hand, the decoder (100) may be configured to use a time constant that is relatively long compared to the duration of the auditory event in response to a gain change at a release (e.g., a decrease in loudness level) in the auditory event. Thus, the gain change at the release can be implemented relatively slowly by the decoder (100), whereby sounds that are supposed to be perceived as constant or gradually decaying may not be auditorily or perceptually disrupted. The quick response at the attack at the auditory event boundary and the slow response at the release in the auditory event allow for a fast perception of the arrival of the auditory event and preserve the perceptual quality and / or integrity between auditory events such as piano chords, which include loud and soft sounds linked by specific loudness level relationships and / or specific time relationships.

[0138] In some embodiments, the auditory event boundaries indicated by the auditory events and ASA parameters are used by the decoder (100) to control gain changes in one, two, some, or all of the channels in a particular speaker configuration in the decoder (100).

[0139] 〈15. Loudness Level Transition〉 Loudness level transitions can occur, for example, between two programs, between a program and a loud commercial, etc. In some embodiments, the decoder (100) is configured to maintain a histogram of the instantaneous loudness levels based on past audio content (such as received from the audio signal 102 encoded over the past 4 seconds, for example). Over a time interval from before the loudness level transition to after the loudness level transition, two regions with an increased probability can be recorded in the histogram. One of those regions is centered around the previous loudness level, while the other of those regions is centered around the new loudness level.

[0140] The decoder (100) may dynamically determine a smoothed loudness level when the audio content is processed and determine the corresponding bin in the histogram (e.g., the bin of the instantaneous loudness level that includes the same value as the smoothed loudness level) based on the smoothed loudness level. The decoder (100) is further configured to compare the probability in the corresponding bin with a threshold (e.g., 6%, 7%, 7.5%, etc.). Here, the total area under the histogram curve (e.g., the sum of all bins) represents a probability of 100%. The decoder can be configured to detect the occurrence of a loudness level transition by determining that the probability of the corresponding bin is below the threshold. In response, the decoder (100) may be configured to select a relatively small time constant to adapt relatively quickly to the new loudness level. As a result, the duration of the loud (or soft) start within the loudness level transition can be shortened.

[0141] In some embodiments, the decoder (100) uses a silence / noise gate to prevent low instantaneous loudness levels from falling into the histogram and becoming high probability bins in the histogram. Additionally, optionally, or alternatively, the decoder (100) may be configured to use the ASA parameters to detect auditory events to be included in the histogram. In some embodiments, the decoder (100) determines the time-dependent value of the time-averaged auditory event certainty metric

Number

Number

Number

[0142] In some embodiments, for the loudness levels (e.g., instantaneous, etc.) that are allowed to be included in the histogram (e.g., the value of the corresponding A[t] with a bar is above the histogram inclusion threshold), weights are assigned such that the loudness level is the same as, or proportional to, the time-dependent value of the time-averaged auditory event certainty metric [A[t] with a bar] that is contemporaneous with those loudness levels. As a result, loudness levels closer to the auditory event boundary have more influence on the histogram than other loudness levels that are not close to the auditory event boundary (e.g., the A[t] with a bar has a relatively large value, etc.).

[0143] 〈16. Reset〉 In some embodiments, the encoder (e.g., 150, etc.) described in this document is configured to detect a reset event and include an indicator of the reset event in the encoded audio signal (e.g., 102, etc.). In a first example, the encoder (150) detects a reset event in response to determining that a continuous period of relative silence (e.g., 250 milliseconds, etc., configurable by the system and / or user) has occurred. In a second example, the encoder (150) detects a reset event in response to determining that a large instantaneous drop in excitation level occurs across all frequency bands. In a third example, the encoder is provided with an input (e.g., user input, system-controlled metadata, etc.) where a content transition (e.g., program start / end, scene change, etc.) that requests a reset occurs.

[0144] In some embodiments, the decoder (e.g., 100, etc.) described in this document implements a reset mechanism that can be used to speed up instantaneous gain smoothing. The reset mechanism is useful and may be invoked when a switch occurs between channels and audiovisual inputs.

[0145] In some embodiments, the decoder (100) may be configured to determine whether a reset event occurs by determining whether a period of relative silence occurs (e.g., 250 milliseconds configurable by the system and / or user), or whether a large instantaneous drop in the excitation level occurs across all frequency bands.

[0146] In some embodiments, the decoder (100) is configured to discriminate that a reset event occurs in response to receiving an indicator (such as a reset event) provided in the encoded audio signal (102) by an upstream encoder (such as 150).

[0147] The reset mechanism may be adapted to issue a reset when the decoder (100) determines that a reset event has occurred. In some embodiments, the reset mechanism is configured to use a more aggressive cut behavior of the DRC compression curve to prevent a hard start (such as a loud program / channel / audio-visual source). Additionally, optionally, or alternatively, the decoder (100) may be configured to implement a safeguard to gracefully recover when the decoder (100) detects that the reset has been accidentally triggered.

[0148] 〈17. Gain Provided by Encoder〉 In some embodiments, the audio decoder can be configured to calculate one or more sets of gains (e.g., DRC gains, etc.) for individual portions of the audio content to be encoded in the encoded audio signal (e.g., audio data blocks, audio data frames, etc.). Those sets of gains generated by the audio encoder can include a first set of gains including a single broadband (or wideband) gain for all channels (e.g., left front, right front, low frequency effect or LFE, center, left surround, right surround, etc.); a second set of gains including individual broadband (or wideband) gains for individual subsets of channels; a third set of gains including individual broadband (or wideband) gains for individual subsets of channels and for each of a first number (e.g., two, etc.) of individual bands (e.g., two bands in each channel, etc.); a fourth set of gains including individual broadband (or wideband) gains for individual subsets of channels and for each of a second number (e.g., four, etc.) of individual bands (e.g., four bands in each channel, etc.); and so on. The subsets of channels described herein can be one of a subset including the left front, right front, and LFE channels, a subset including the center channel, a subset including the left surround and right surround channels, etc.

[0149] In some embodiments, an audio encoder is configured to transmit, in a time-synchronized manner, one or more portions of audio content (e.g., audio data blocks, audio data frames, etc.) and one or more sets of gains calculated for the one or more portions of audio content. An audio decoder that receives the one or more portions of audio content can select and apply a set of gains, out of the one or more sets of gains, with little or no latency. In some embodiments, the audio encoder can implement a subframing technique in which the one or more sets of gains are carried in one or more subframes as shown in FIG. 4 (e.g., using differential coding, etc.). In one example, the subframes may be encoded within the audio data block or audio data frame for which their gains are calculated. In another example, the subframes may be encoded within audio data blocks or audio data frames preceding the audio data block or audio data frame for which their gains are calculated. In another non-limiting example, the subframes may be encoded within audio data blocks or audio data frames within a time period from the audio data block or audio data frame for which their gains are calculated. In some embodiments, Huffman and differential coding may be used to load data into and / or compress the subframes carrying those sets of gains.

[0150] 〈18. Exemplary Systems and Process Flows〉 FIG. 5 shows an exemplary codec system in a non-limiting exemplary embodiment. A content creator, which may be a processing unit within an audio encoder such as 150, is configured to provide audio content (“audio”) to an encoder unit (“NGC encoder”). The encoder unit formats the audio content into audio data blocks and / or frames and encodes the audio data blocks and / or frames into an encoded audio signal. The content creator is also configured to establish / generate one or more programs in the audio content, one or more dialog loudness levels (“dialnorm”) such as commercials, etc., and one or more dynamic range compression curve identifiers (“compression curve ID”). The content creator may determine the dialog loudness level from one or more dialog audio tracks in the audio content. The dynamic range compression curve identifier may be selected based at least in part on user input, system configuration setting parameters, etc. The content creator may be a human (e.g., an artist, an audio engineer, etc.) who uses tools to generate the audio content and dialnorm.

[0151] Based on the dynamic range compression curve identifier, the encoder (150) generates one or more sets of DRC parameters including, but not limited to, corresponding reference dialog loudness levels ("reference levels") for a plurality of playback environments supported by the one or more dynamic range compression curves. These sets of DRC parameters may be encoded in the metadata of the encoded audio signal, in-band with the audio content, out-of-band with the audio content, etc. Operations such as compression, format multiplexing ("MUX"), etc. may be performed as part of generating an encoded audio signal that can be delivered to an audio decoder such as 100. The encoded audio signal may be encoded with a syntax that supports carrying audio data elements, sets of DRC parameters, reference loudness levels, dynamic range compression curves, functions, look-up tables, Huffman codes used in compression, sub-frames, etc. In some embodiments, the syntax allows an upstream device (e.g., an encoder, decoder, transcoder, etc.) to transmit a gain to a downstream device (e.g., a decoder, transcoder, etc.). In some embodiments, the syntax used to encode data into and / or decode data from the encoded audio signal is configured to support backward compatibility such that a device that relies on the gain calculated by an upstream device may optionally continue to do so.

[0152] In some embodiments, the encoder (150) calculates two or more sets of gains for the audio content (such as gain smoothing using an appropriate reference dialog loudness level, DRC gain, etc.). These sets of gains may be provided in the metadata encoded in the audio signal encoded together with the audio content to provide the one or more dynamic range compression curves. The first set of gains may correspond to the broadband (or wideband) gain for all channels in a speaker configuration or profile (such as default). The second set of gains may correspond to the broadband (or wideband) gain for each of all channels in a speaker configuration or profile. The third set of gains may correspond to the broadband (or wideband) gain for each of two bands in each of all channels in a speaker configuration or profile. The fourth set of gains may correspond to the broadband (or wideband) gain for each of four bands in each of all channels in a speaker configuration or profile. In some embodiments, a set of gains calculated for a particular speaker configuration may be transmitted in the metadata together with the dynamic range compression curve for that speaker configuration (such as parameterized). In some embodiments, a set of gains calculated for a particular speaker configuration may replace the dynamic range compression curve for that speaker configuration (such as parameterized) in the metadata. Additional speaker configurations or profiles may be supported under the techniques described herein.

[0153] Decoder (100) is configured to extract an audio data block and / or frame and metadata from the encoded audio signal through operations such as decompression, format conversion, and demultiplexing ( "DEMUX"). The extracted audio data block and / or frame may be decoded into audio data elements or samples by a decoder unit ( "NGC decoder"). Decoder (100) is further configured to determine a profile for a specific playback environment in which the audio content is to be rendered in decoder (100) and select a dynamic range compression curve from the metadata extracted from the encoded audio signal. Digital audio processing unit ( "DAP") is configured to apply DRC or other operations to audio data elements or samples for the purpose of generating an audio signal that drives an audio channel in a specific playback environment. Decoder (100) can calculate and apply a DRC gain based on the audio data block or frame and the selected dynamic range compression curve. Decoder (100) can also adjust the output dialog loudness level based on the reference dialog loudness level associated with the selected dynamic range compression curve and the dialog loudness level in the metadata extracted from the encoded audio signal. Decoder (100) can then apply a gain limiter specific to the playback scenario related to the audio content and the specific playback environment. In this way, decoder (100) can render / play the audio content as adjusted according to the playback scenario.

[0154] Figure 5A shows another exemplary decoder (which may be the same as decoder 100 of FIG. 5). As shown in FIG. 5A, the decoder of FIG. 5A is configured to extract audio data blocks and / or frames and metadata from an encoded audio signal through operations such as decompression, format removal, multiplex separation ("DEMUX"). The extracted audio data blocks and / or frames may be decoded into audio data elements or samples by a decoder unit ("decode"). The decoder of FIG. 5A is further configured to perform DRC gain calculations based on a default compression curve for a set of default gains, a smoothing constant related to the default compression curve, etc. The decoder of FIG. 5A is further configured to extract a set of differential gains for a non-default gain profile from profile-related metadata in the metadata, determine a set of non-differential gains for the non-default gain profile in the decoder of FIG. 5A where the audio content is rendered, and apply the set of non-differential gains and other operations to the audio data elements or samples to generate a DRC-improved audio output that drives the audio channels in a particular playback environment. The decoder of FIG. 5A can render / play the audio content according to the non-default gain profile, whether or not the decoder of FIG. 5A itself implements support for obtaining a set of non-differential gains directly for the non-default gain profile by performing a set of gain generation operations.

[0155] FIGS. 6A through 6D show an exemplary process flow. In some embodiments, one or more computing devices or units in a media processing system may execute this process flow.

[0156] Figure 6A shows an exemplary process flow that may be implemented by the audio decoder described in this paper. In block 602 of Figure 6A, a first device (e.g., the audio decoder 100 of Figure 1A, etc.) receives an audio signal including audio content and definition data about one or more dynamic range compression curves.

[0157] In block 604, the first device determines a specific playback environment.

[0158] In block 606, the first device establishes a specific dynamic range compression curve for the specific playback environment based on the definition data about the one or more dynamic range compression curves extracted from the audio signal.

[0159] In block 608, the first device performs one or more dynamic range control (DRC) operations on one or more portions of the audio content extracted from the audio signal. The one or more DRC operations are at least partially based on one or more DRC gains obtained from the specific dynamic range compression curve.

[0160] In some embodiments, the definition data about the one or more dynamic range compression curves includes an attack time, a release time, or a reference loudness level related to at least one of the one or more dynamic range compression curves.

[0161] In some embodiments, the first device is further configured to perform steps such as calculating one or more loudness levels for the one or more portions of the audio content; determining the one or more DRC gains based on the specific dynamic range compression curve and the one or more loudness levels for the one or more portions of the audio content.

[0162] In one embodiment, at least one of the loudness levels calculated for the one or more portions of the audio content is one or more of a specific loudness level related to one or more frequency bands, a broadband loudness level that traverses a broadband range, a wideband loudness level that traverses a wideband range, a broadband loudness level that traverses a plurality of frequency bands, a wideband loudness level that traverses a plurality of frequency bands, and the like.

[0163] In one embodiment, at least one of the loudness levels calculated for the one or more portions of the audio content is one or more of an instantaneous loudness level or a loudness level smoothed over one or more time intervals.

[0164] In one embodiment, the one or more operations are related to one or more of adjusting a dialogue loudness level, gain smoothing, gain limiting, dynamic equalization, noise compensation, and the like.

[0165] In one embodiment, the first device is further configured to: extract one or more dialogue loudness levels from the encoded audio signal; adjust the one or more dialogue loudness levels to one or more reference dialogue loudness levels; and the like.

[0166] In one embodiment, the first device is further configured to: extract one or more auditory scene analysis (ASA) parameters from the encoded audio signal; vary one or more time constants used in smoothing the gain applied to the audio content, where the gain is related to one or more of the one or more DRC gains; and perform gain smoothing or gain limiting, and the like.

[0167] In one embodiment, the first device further: determines that a reset event occurs in the one or more portions of the audio content based on an indicator of the reset event, wherein the indicator of the reset is extracted from the encoded audio signal; in response to determining that the reset event occurs in the one or more portions of the audio content, performs one or more actions on one or more gain smoothing operations being performed at the time of determining that the reset event occurs in the one or more portions of the audio content; etc., and is configured to perform.

[0168] In one embodiment, the first device further: maintains a histogram of instantaneous loudness levels, wherein the histogram contains instantaneous loudness levels calculated from a time interval in the audio content; determines whether a specific loudness level is above a threshold in a high probability region of the histogram, wherein the specific loudness level is calculated from a portion of the audio content; in response to determining that the specific loudness level is above the threshold in the high probability region of the histogram, determines that a loudness transition is occurring and shortens a time constant used in gain smoothing to speed up the loudness transition; etc., and is configured to perform.

[0169] FIG. 6B shows an exemplary process flow that may be implemented by the audio encoder described in this document. In block 652 of FIG. 6B, a second device (such as the audio encoder 150 of FIG. 1B, etc.) receives audio content in a source audio format.

[0170] In block 654, the second device obtains definition data for one or more dynamic range compression curves.

[0171] In block 656, the second device generates an audio signal including the audio content and the definition data for the one or more dynamic range compression curves.

[0172] In certain embodiments, the second device is further configured to: determine one or more identifiers for the one or more dynamic range compression curves; retrieve the definition data for the one or more dynamic range compression curves from a reference data storage based on the one or more identifiers; and so on.

[0173] In certain embodiments, the second device is further configured to: calculate one or more dialogue loudness levels for the one or more portions of the audio content; encode the one or more dialogue loudness levels together with the one or more portions of the audio content into the encoded audio signal; and so on.

[0174] In certain embodiments, the second device is configured to: perform an auditory scene analysis (ASA) on the one or more portions of the audio content; generate one or more ASA parameters based on the results of the ASA for the one or more portions of the audio content; encode the one or more ASA parameters together with the one or more portions of the audio content into the encoded audio signal; and so on.

[0175] In certain embodiments, the second device is further configured to: determine that one or more reset events occur in the one or more portions of the audio content; encode one or more indicators of the one or more reset events together with the one or more portions of the audio content into the encoded audio signal; and so on.

[0176] In one embodiment, the second device is further configured to encode the one or more portions of the audio content into one or more of an audio data frame or an audio data block.

[0177] In one embodiment, the first DRC gain of the one or more DRC gains applies to each channel in a first proper subset in a set of all channels in a particular speaker configuration corresponding to that particular playback environment, while a second different DRC gain of the one or more DRC gains applies to each channel in a second proper subset in the set of all channels in the particular speaker configuration corresponding to that particular playback environment.

[0178] In one embodiment, the first DRC gain of the one or more DRC gains applies to a first frequency band, and a second different DRC gain of the one or more DRC gains applies to a second different frequency band.

[0179] In one embodiment, the one or more portions of the audio content include one or more of an audio data frame or an audio data block. In one embodiment, the encoded audio signal is part of an audiovisual signal.

[0180] In one embodiment, the one or more DRC gains are defined in a loudness region.

[0181] FIG. 6C shows an exemplary process flow that may be implemented by the audio decoder described in this document. In block 662 of FIG. 6C, a third device (e.g., the audio decoder 100 of FIG. 1A, the audio decoder of FIG. 5, the audio decoder of FIG. 5A, etc.) receives an audio signal that includes audio content, definition data for one or more dynamic range compression (DRC) curves, and one or more sets of differential gains.

[0182] In block 664, the third device identifies a particular set of differential gains for a gain profile in a particular playback environment from among the one or more sets of differential gains. The third device also identifies a default DRC curve related to the particular set of differential gains from among the one or more DRC curves.

[0183] In block 666, the third device generates a set of default gains, at least in part, based on the default DRC curve.

[0184] In block 668, based at least in part on a combination of the set of default gains and the particular set of differential gains, the third device performs one or more operations on one or more portions of the audio content extracted from the audio signal.

[0185] In some embodiments, the set of default gains includes non-differential gains generated by performing a set of gain generation operations, at least in part, based on the default DRC curve.

[0186] In some embodiments, the default DRC curve represents a default gain profile. In some embodiments, the particular set of differential gains in relation to the default DRC curve represents a non-default gain profile. In some embodiments, the audio signal does not include definition data for a non-default DRC curve corresponding to the non-default gain profile.

[0187] In one embodiment, the particular set of differential gains includes the gain difference between a set of non-differential non-default gains generated for a non-default gain profile and a set of non-differential default gains generated for the default gain profile represented by the default DRC curve. The set of non-differential non-default gains and the set of non-differential default gains may be generated by an upstream audio decoder that encodes the audio signal.

[0188] In one embodiment, at least one set of the non-differential non-default gains or the set of non-differential default gains is not provided as part of the audio signal.

[0189] FIG. 6D shows an exemplary process flow that may be implemented by the audio decoder described herein. At block 672 of FIG. 6D, a fourth device (e.g., audio encoder 150 of FIG. 1A, audio encoder of FIG. 5, etc.) receives audio content in a source audio format.

[0190] At block 674, the fourth device generates a set of default gains based at least in part on a default dynamic range compression (DRC) curve representing a default gain profile.

[0191] At block 676, the fourth device generates a set of non-default gains for a non-default gain profile.

[0192] At block 678, based at least in part on the set of default gains and the set of non-default gains, the fourth device generates a set of differential gains. The set of differential gains represents the non-default gain profile in relation to the default DRC curve.

[0193] In block 680, a fourth apparatus generates an audio signal that includes the audio content and the defined data for one or more DRC curves and for one or more sets of differential gains. The one or more sets of differential gains include the set of differential gains.

[0194] In some embodiments, the non-default gain profile is represented by a DRC curve. In one embodiment, the audio signal does not include defined data for the DRC curve representing the non-default gain profile. In some embodiments, the non-default gain profile is not represented by a DRC curve.

[0195] In one embodiment, an apparatus having a processor and configured to execute any of the methods described herein.

[0196] In one embodiment, a non-transitory computer-readable storage medium that includes software instructions that, when executed by one or more processors, cause execution of any of the methods described herein. Note that although multiple separate embodiments are discussed herein, any combination of the embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.

[0197] 〈19. Implementation Mechanism - Overview of Hardware〉 According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be fixedly configured to execute the present techniques, or may include one or more digital electronic devices that are persistently programmed to execute the present techniques, such as one or more application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to execute the present techniques according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices may achieve the present techniques by combining custom fixed configuration logic, ASICs, or FPGAs with custom programming. The special-purpose computing devices may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device incorporating fixed configuration and / or program logic for implementing the present techniques.

[0198] For example, FIG. 7 is a block diagram showing a computer system 700 in which an embodiment of the present invention may be implemented. The computer system 700 includes a bus 702 or other communication mechanism for communicating information, and a hardware processor 704 coupled to the bus 702 for processing information. The hardware processor 704 may be, for example, a general-purpose microprocessor.

[0199] Computer system 700 also includes a main memory 706 coupled to a bus 702 for storing information and instructions to be executed by a processor 704, such as random access memory (RAM) or other dynamic storage device. The main memory 706 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 704. When such instructions are stored on a non-transitory storage medium accessible to the processor 704, the computer system 700 is caused to be a special purpose machine specific to the device to execute the processing specified in the instructions.

[0200] Computer system 700 further includes a read only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions for the processor 704. A storage device 710, such as a magnetic disk or optical disk, is provided and coupled to the bus 702 for storing information and instructions.

[0201] Computer system 700 may be coupled via the bus 702 to a display 712, such as a liquid crystal display (LCD), for displaying information to a computer user. An input device 714 including alphanumeric and other keys is coupled to the bus 702 for communicating information and command selections to the processor 704. Another type of user input device is a cursor control 716, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 704 and for controlling cursor movement on the display 712. This input device typically has two degrees of freedom in two axial directions of a first axis (e.g., x) and a second axis (e.g., y), whereby the device can specify a position within a plane.

[0202] Computer system 700 may use device-specific fixed configuration logic, one or more ASICs or FPGAs, firmware and / or program logic that, in combination with the computer system, configure or program the computer system 700 as a special purpose machine to implement the techniques described herein. According to one embodiment, the techniques herein are performed by computer system 700 in response to one or more sequences of one or more instructions included in main memory 706 being executed by processor 704. Such instructions may be read into main memory 706 from another storage medium such as storage device 710. Execution of the sequence of instructions included in main memory 706 causes processor 704 to perform the process steps described herein. In an alternative embodiment, fixed configuration circuitry may be used instead of or in combination with software instructions.

[0203] As used herein, the term “storage medium” refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 710. Volatile media includes dynamic memory, such as main memory 706. Common forms of storage media include, for example, floppy disks, flexible disks, hard disk drives, semiconductor drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM, and EPROM, flash EPROM, NVRAM, any other memory chip or cartridge.

[0204] A memory medium is different from, but may be used in association with, a transmission medium. The transmission medium participates in transferring information between memory media. For example, the transmission medium includes coaxial cables, copper wire, and fiber optics, including the wires that form bus 702. The transmission medium can also take the form of acoustic or light waves such as those generated during radio and infrared data communications.

[0205] Various forms of media can be involved in carrying one or more sequences of one or more instructions to processor 704 for execution. For example, the instructions may initially be carried on a magnetic disk or semiconductor drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions through a telephone line using a modem. A modem local to computer system 700 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on bus 702. Bus 702 carries the data to main memory 706, from which processor 704 fetches and executes the instructions. The instructions received by main memory 706 may optionally be stored on storage device 710 before or after execution by processor 704.

[0206] Computer system 700 also includes a communication interface 718 coupled to bus 702. The communication interface 718 provides a two-way data communication coupling to network link 720 that is connected to local network 722. For example, the communication interface 718 may be an integrated services digital communication network (ISDN) card, cable modem, satellite modem, or a modem for providing a data communication connection to a corresponding type of telephone line. As another example, the communication interface 718 may be a local area network (LAN) card for providing a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, the communication interface 718 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0207] Network link 720 typically provides data communication to other data devices through one or more networks. For example, network link 720 may provide a connection through local network 722 to data facilities operated by host computer 724 or an Internet service provider (ISP) 726. The ISP 726 provides data communication services through a worldwide packet data communication network commonly referred to as the "Internet" 728 today. Both the local network 722 and the Internet 728 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks that carry digital data to / from the computer system 700, and signals on network link 720 and through communication interface 718 are exemplary forms of transmission media.

[0208] Computer system 700 can send messages and receive data, including program code, through a network (s), network link 720, and communication interface 718. In the example of the Internet, server 730 may send the requested code for an application program through Internet 728, ISP 726, local network 722, and communication interface 718.

[0209] The received code may be executed by processor 704 upon receipt and / or stored in storage device 710 or other non-volatile storage for later execution.

[0210] 〈20. Equivalents, Extensions, Alternatives and Others〉 In the above specification, exemplary embodiments of the invention have been described while referring to numerous individual details that may vary depending on the implementation. Thus, the only and exclusive indicator of what the invention is and what is intended by the applicant to be the invention is the claims of the patent granted for this application, including any subsequent amendments, in the specific form in which such claims are patented. If there are definitions explicitly set forth in this document for the terms included in such claims, then those definitions govern the meaning of the terms used in the claims. Therefore, limitations, elements, attributes, features, advantages, or characteristics not explicitly recited in the claims should not in any way limit the scope of such claims. Thus, the specification and drawings should be regarded in an illustrative rather than a restrictive sense.

[0211] Some aspects will be described. [Aspect 1] Receiving an audio signal including audio content and one or more sets of differential gains; Identifying a particular set of differential gains for a gain profile in a particular playback environment among the one or more sets of differential gains; generating a set of default gains based at least on the default dynamic range compression (DRC) curve associated with the specific set of differential gains; performing one or more operations on one or more portions of the audio content extracted from the audio signal, based at least in part on a combination of the set of default gains and the specific set of differential gains; A method performed by one or more computers. [Aspect 2] The method according to aspect 1, wherein the set of default gains includes non-differential gains generated by performing a set of gain generation operations based at least in part on the default DRC curve. [Aspect 3] The method according to aspect 1 or 2, wherein the default DRC curve represents a default gain profile. [Aspect 4] The method according to any one of aspects 1 to 3, wherein the specific set of differential gains in relation to the default DRC curve represents a non-default gain profile. [Aspect 5] The method according to aspect 4, wherein the audio signal does not include definition data for a non-default DRC curve corresponding to the non-default gain profile. [Aspect 6] The method according to any one of aspects 1 to 5, wherein the specific set of differential gains includes the gain difference between the set of non-differential non-default gains generated for the non-default gain profile and the set of non-differential default gains generated for the default gain profile represented by the default DRC curve. [Aspect 7] The method according to aspect 6, wherein the set of non-differential non-default gains and the set of non-differential default gains are generated by an upstream audio decoder that encodes the audio signal. [Aspect 8] The method according to aspect 6, wherein at least one of the set of non-differential non-default gains or the set of non-differential default gains is not provided as part of the audio signal. [Aspect 9] The method according to any one of aspects 1 to 8, wherein the definition data for the one or more DRC curves includes one or more of an attack time, a release time, or a reference loudness level related to at least one of the one or more DRC curves. [Aspect 10] The method according to aspect 9, wherein the reference loudness level represents a target range of a playback level for rendering the audio content by an audio decoder. [Aspect 11] Calculating one or more loudness levels for the one or more portions of the audio content; Generating a set of non-differential non-default gains based on the set of non-differential default gains and the specific set of differential gains; Further comprising applying the set of non-differential non-default gains to the one or more portions of the audio content. The method according to any one of aspects 1 to 10. [Aspect 12] The method according to aspect 11, wherein at least one of the one or more loudness levels calculated for the one or more portions of the audio content is one or more of a specific loudness level related to one or more frequency bands, a broadband loudness level over a broadband range, a wideband loudness level over a wideband range, a broadband loudness level over a plurality of frequency ranges, or a wideband loudness level over a plurality of frequency ranges. [Aspect 13] The method according to aspect 11, wherein at least one of the one or more loudness levels calculated for the one or more portions of the audio content is an instantaneous loudness level or one or more smoothed loudness levels over one or more time intervals. 〔Aspect 14〕 The method according to any one of aspects 1 to 13, wherein the one or more operations include one or more operations related to adjusting a dialog loudness level, gain smoothing, gain limiting, dynamic equalization, or noise compensation. 〔Aspect 15〕 The method according to any one of aspects 1 to 14, wherein the method is executed by an audio decoding device, and the default DRC curve is defined in the audio decoding device. 〔Aspect 16〕 Receiving definition data for one or more dynamic range compression (DRC) curves; Further comprising identifying a default DRC curve related to the specific set of differential gains among the one or more DRC curves. The method according to any one of aspects 1 to 15. 〔Aspect 17〕 Extracting one or more auditory scene analysis (ASA) parameters from the encoded audio signal; Further comprising varying one or more time constants used in smoothing the gain applied to the audio content. The method according to any one of aspects 1 to 16. 〔Aspect 18〕 Determining that a reset event occurs in the one or more portions of the audio content based on an indicator of the reset event, wherein the indicator of the reset is extracted from the encoded audio signal. In response to determining that the reset event occurs in the one or more portions of the audio content, performing one or more actions on one or more gain smoothing operations being performed at the time when it is determined that the reset event occurs in the one or more portions of the audio content. The method according to any one of aspects 1 to 17. [Aspect 19] The method according to aspect 18, wherein at least one of the one or more smoothing operations uses a first smoothing time constant before the reset event, and at least one of the one or more smoothing operations uses a second smoothing time constant smaller than the first smoothing time constant in response to determining that the reset event occurs. [Aspect 20] Maintaining a histogram of instantaneous loudness levels, the histogram being filled with instantaneous loudness levels calculated from a time interval in the audio content; Determining whether a specific loudness level is below a threshold in a high probability region of the histogram, the specific loudness level being calculated from a part of the audio content; In response to determining that the specific loudness level is below the threshold in the high probability region of the histogram: Determining that a loudness transition is occurring, Shortening a time constant used in gain smoothing to speed up the loudness transition. The method according to any one of aspects 1 to 19. [Aspect 21] The specific set of differential gains includes first differential gains associated with each channel in a first proper subset of the set of all channels in a particular speaker configuration, and the specific set of differential gains includes second differential gains associated with each channel in a second proper subset of the set of all channels in the particular speaker configuration, the method according to any one of aspects 1 to 20. [Aspect 22] The specific set of differential gains includes first differential gains associated with a first frequency band, and the specific set of differential gains includes second different differential gains associated with a second different frequency band, the method according to any one of aspects 1 to 21. [Aspect 23] The one or more portions of the audio content include one or more of audio data frames, audio data blocks, or audio samples, the method according to any one of aspects 1 to 22. [Aspect 24] The specific set of differential gains is defined in a loudness region, the method according to any one of aspects 1 to 23. [Aspect 25] The encoded audio signal is part of an audiovisual signal, the method according to any one of aspects 1 to 24. [Aspect 26] Receiving audio content in a source audio format; Generating a set of default gains, at least in part based on a default dynamic range compression (DRC) curve, wherein the default DRC curve represents a default gain profile; Generating a set of non-default gains for a non-default gain profile; Generating the set of differential gains, at least in part based on the set of default gains and the set of non-default gains, wherein the set of differential gains represents the non-default gain profile in relation to the default DRC curve; generating an audio signal including the audio content and the one or more sets of differential gains including the set of differential gains A method executed by one or more computing devices. 〔Aspect 27〕 The method according to aspect 26, wherein the non-default gain profile is represented by a DRC curve. 〔Aspect 28〕 The method according to aspect 27, wherein the audio signal does not include definition data for the DRC curve representing the non-default gain profile. 〔Aspect 29〕 The method according to any one of aspects 26 to 28, wherein the non-default gain profile is not represented by a DRC curve. 〔Aspect 30〕 determining one or more identifiers for the one or more dynamic range compression curves; further comprising retrieving the definition data for the one or more dynamic range compression curves from a reference data storage based on the one or more identifiers. The method according to any one of aspects 26 to 29. 〔Aspect 31〕 The method according to any one of aspects 26 to 30, wherein the set of default gains includes first non-differential gains generated by performing a first set of gain generation operations based at least in part on the default DRC curve, and the set of non-default gains includes second non-differential gains generated by performing a second set of gain generation operations for the non-default gain profile. 〔Aspect 32〕 calculating one or more dialogue loudness levels for one or more portions of the audio content; further comprising encoding the one or more dialogue loudness levels together with the one or more portions of the audio content into the encoded audio signal. The method according to any one of aspects 26 to 31. [Aspect 33] The method according to aspect 32, wherein at least one of the one or more dialog loudness levels is determined from one or more audio tracks including dialog audio content. [Aspect 34] Performing auditory scene analysis (ASA) on the one or more portions of the audio content; Generating one or more ASA parameters based on the results of the ASA on the one or more portions of the audio content; Further comprising encoding the one or more ASA parameters together with the one or more portions of the audio content into the encoded audio signal. The method according to any one of aspects 26 to 33. [Aspect 35] Determining that one or more reset events occur in one or more portions of the audio content; Further comprising encoding one or more indicators of the one or more reset events together with the one or more portions of the audio content into the encoded audio signal. The method according to any one of aspects 26 to 34. [Aspect 36] The method according to any one of aspects 26 to 35, further comprising encoding the one or more portions of the audio content into one or more audio data frames or audio data blocks. [Aspect 37] The method according to any one of aspects 26 to 36, wherein at least one of the one or more dynamic range compression curves is defined in a loudness region. [Aspect 38] The method according to any one of aspects 26 to 37, wherein the encoded audio signal is part of an audiovisual signal. [Aspect 39] The method according to any one of aspects 26 to 38, wherein the definition data for the one or more dynamic range compression curves includes one or more sets of parameters, and at least one set of the one or more sets of parameters represents one or more of a look-up table, a curve, or a plurality of segment piecewise linear lines. [Aspect 40] The method according to any one of aspects 26 to 39, wherein the encoded audio signal includes an indicator for selecting a DRC curve defined in a receiving device as the default DRC curve. [Aspect 41] Sending definition data for various DRC curves in the encoded audio signal; Further including including an indicator for selecting the default DRC curve from among the one or more DRC curves. The method according to any one of aspects 26 to 40. [Aspect 42] A media processing system configured to execute the method according to any one of aspects 1 to 41. [Aspect 43] An apparatus having a processor configured to execute the method according to any one of aspects 1 to 41. [Aspect 44] A non-transitory computer-readable storage medium including software instructions that, when executed by one or more processors, cause execution of the method according to any one of aspects 1 to 41.

Claims

1. A method for dynamic range control (DRC) of an audio signal, the method comprising: receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, the audio signal including audio sample data for each channel of the reference channel configuration and DRC metadata generated by an encoder, the DRC metadata generated by the encoder including DRC gains for a plurality of channel configurations, the DRC gains for the plurality of channel configurations including a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; downmixing the audio sample data to downmixed audio sample data for the channels of the playback channel configuration; selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the downmixed audio sample data to generate output audio sample data for each channel of the playback channel configuration; wherein the audio signal is framed into frames, each frame including one or more sub-frames, and the set of DRC gains for the playback channel configuration includes one DRC gain per sub-frame; A method.

2. A non-transitory computer-readable storage medium storing software instructions, the software instructions, when executed by one or more processors: Receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, wherein the audio signal includes audio sample data for each channel of the reference channel configuration and dynamic range control (DRC) metadata generated by an encoder, the DRC metadata generated by the encoder includes DRC gains for a plurality of channel configurations, and the DRC gains for the plurality of channel configurations include a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; Downmixing the audio sample data to downmixed audio sample data for channels of the playback channel configuration; Selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; Applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the downmixed audio sample data to generate output audio sample data for each channel of the playback channel configuration, wherein the audio signal is framed into frames, each frame includes one or more sub-frames, and the set of DRC gains for the playback channel configuration includes one DRC gain per sub-frame. Non-transitory computer-readable storage medium. Claim 3 An audio signal processing apparatus for dynamic range control of an audio signal, the audio signal processing apparatus comprising: Receiving, by an audio decoder operating in a playback channel configuration different from a reference channel configuration, an audio signal for the reference channel configuration, wherein the audio signal includes audio sample data for each channel of the reference channel configuration and dynamic range control (DRC) metadata generated by an encoder, the DRC metadata generated by the encoder includes DRC gains for a plurality of channel configurations, and the DRC gains for the plurality of channel configurations include a set of DRC gains for the playback channel configuration and a set of DRC gains for the reference channel configuration; downmixing the audio sample data to produce downmixed audio sample data for the channels of the playback channel configuration; selecting the set of DRC gains for the playback channel configuration from the DRC gains for the plurality of channel configurations; applying the set of DRC gains for the playback channel configuration as part of an overall gain applied to the downmixed audio sample data to produce output audio sample data for each channel of the playback channel configuration, wherein the audio signal is organized into frames, each frame including one or more sub-frames, and the set of DRC gains for the playback channel configuration includes one DRC gain per sub-frame, an apparatus. **Claim 4** A computer program product having executable instructions for performing the method of claim 1 when executed on a computer.

Citation Information

Patent Citations

  • Method and apparatus for encoding and decoding object-based audio signals.

    JP2010508545A

  • Sound signal conversion device and sound signal conversion program

    JP2012034295A

  • Audio metadata transcoding

    JP2012504260A

  • Protecting signal clipping using existing audio gain metadata

    JP2012507059A

  • Audio decoder and decoding method using efficient downmixing

    JP2012527021A