Efficient DRC profile transmission

By encoding multiple DRC profiles within audio frames and enabling dynamic selection based on rendering modes, the system addresses the challenge of maintaining audio quality and intelligibility across varied playback systems, ensuring effective dynamic range control and reduced distortion.

JP2026086745APending Publication Date: 2026-05-26DOLBY INTERNATIONAL AB

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2026-02-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing media playback systems face challenges in maintaining high-quality and intelligible audio reproduction across a wide range of rendering devices with different capabilities and environments, due to varying dynamic range requirements.

Method used

A method and system for encoding and decoding audio signals that include multiple dynamic range control (DRC) profiles within frames, allowing decoders to select appropriate DRC profiles for different rendering modes, ensuring high-quality and intelligible audio reproduction.

Benefits of technology

Ensures high-quality and intelligible audio reproduction across diverse playback environments by dynamically controlling the dynamic range, reducing clipping and distortion, and adapting to the specific capabilities of each playback device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086745000001_ABST
    Figure 2026086745000001_ABST
Patent Text Reader

Abstract

It provides efficient dynamic range control (DRC) profile transmission. [Solution] The encoded audio signal 102 has a sequence of frames and represents multiple different DRC profiles for a plurality of different rendering modes. The method includes the steps of: determining a first rendering mode from a plurality of different rendering modes; determining one or more DRC profiles from a subset of DRC profiles contained in the current frame of the sequence of frames; determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode; selecting a default DRC profile if none of the one or more DRC profiles are applicable to the first rendering mode; and decoding the current frame using the current DRC profile.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 058,228, filed on October 1, 2014. The contents of that application are hereby incorporated by reference in their entirety.

[0002] Technical Field This document relates to the processing of audio signals. In particular, this document relates to methods and corresponding systems for transmitting dynamic range control (DRC) profiles in a bandwidth - efficient manner.

Background Art

[0003] The increasing popularity of media consumption devices has created new opportunities and challenges for the creators and distributors of media content for playback on such devices, as well as for the designers and manufacturers of such devices. Many consumer devices can play a wide variety of media content types and formats, including those related to often high - quality, wide - bandwidth, and wide - dynamic - range audio content for HDTV, Blu - ray, or DVD. Media processing devices can be used to play this type of audio content on their internal acoustic transducers or on external transducers such as headphones or high - quality home theater systems. However, all of these playback systems and environments impose significantly different requirements on the dynamic range of the audio signal, either due to various noise levels in the environment or due to the limited capabilities of the playback system to reproduce the required sound pressure level without distortion. Limiting the dynamic range depending on the environment is an approach to provide high quality and intelligibility across a wide range of different rendering devices with different rendering capabilities and listening environments, i.e., across a wide range of rendering modes.

Summary of the Invention

[0004] This paper addresses the technical challenges faced by media content creators and distributors by providing bandwidth-efficient methods for enabling high-quality and intelligible audio signal reproduction across a wide range of rendering devices with different rendering capabilities. [Means for solving the problem]

[0005] In one aspect, a method for generating an encoded audio signal is described. The encoded audio signal has a sequence of frames. The encoded audio signal represents multiple different dynamic range control (DRC) profiles for multiple different rendering modes. The method includes inserting different subsets of DRC profiles from the multiple DRC profiles into different frames of the sequence of frames so that two or more frames of the sequence of frames congruently contain the multiple DRC profiles.

[0006] In a further aspect, a method for decoding an encoded audio signal is described. The encoded audio signal has a sequence of frames. Furthermore, the encoded audio signal represents multiple different Dynamic Range Control (DRC) profiles for multiple different rendering modes. Different subsets of DRC profiles from the multiple DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the multiple DRC profiles. The method includes determining a first rendering mode from the multiple different rendering modes and determining one or more DRC profiles from the subset of DRC profiles contained within the current frame of the sequence of frames. Furthermore, the method includes determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode. Furthermore, if none of the one or more DRC profiles are applicable to the first rendering mode, the method includes selecting a default DRC profile as the current DRC profile. Here, the definition data of the default DRC profile is known in the decoder for decoding the encoded audio signal. Furthermore, this method includes decoding the current frame using the current DRC profile.

[0007] In a further aspect, a bitstream containing an encoded audio signal is described. The encoded audio signal has a sequence of frames. The encoded audio signal represents multiple different Dynamic Range Control (DRC) profiles for multiple different rendering modes. The different subsets of DRC profiles from the multiple DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the multiple DRC profiles.

[0008] Another aspect describes an encoder for generating an encoded audio signal. The encoded audio signal has a sequence of frames. The encoded audio signal represents multiple different Dynamic Range Control (DRC) profiles for multiple different rendering modes. The encoder is configured to insert different subsets of the DRC profiles from the multiple DRC profiles into different frames of the sequence of frames, so that two or more frames of the sequence of frames congruently contain the multiple DRC profiles.

[0009] In a further aspect, a decoder for decoding an encoded audio signal is described. The encoded audio signal has a sequence of frames. The encoded audio signal represents multiple different Dynamic Range Control (DRC) profiles for multiple different rendering modes. The different subsets of DRC profiles from the multiple DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the multiple DRC profiles. The decoder is configured to determine a first rendering mode from the multiple different rendering modes, to determine one or more DRC profiles from the subset of DRC profiles contained within the current frame of the sequence of frames, to determine whether at least one of the one or more DRC profiles is applicable to the first rendering mode, and if none of the one or more DRC profiles are applicable to the first rendering mode, to select a default DRC profile as the current DRC profile. Here, the definition data of the default DRC profile is known to the decoder. The decoder is further configured to decode the current frame using the current DRC profile.

[0010] In a further aspect, a software program is described. This software program may be adapted for execution on a processor and, when executed on a processor, for performing the method steps outlined in this paper.

[0011] Another aspect describes a storage medium. This storage medium may have a software program adapted for execution on a processor and for performing the method steps outlined in this paper when executed on the processor.

[0012] In a further aspect, a computer program product is described. A computer program may have executable instructions for performing the method steps outlined in this paper when executed on a computer.

[0013] It should be noted that methods and systems including preferred embodiments outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any way. In particular, the features of the claims may be combined with each other in any way. [Brief explanation of the drawing]

[0014] The present invention is described below in an illustrative manner with reference to the accompanying drawings. [Figure 1] This is a diagram illustrating an exemplary audio decoder. [Figure 2] This is a diagram illustrating an exemplary audio encoder. [Figure 3] This figure shows an example of a dynamic range compression curve. [Figure 4] This figure shows an example of a dynamic range compression curve. [Figure 5] This figure shows an example sequence of frames. [Figure 6a]This is the first part of a flowchart illustrating an exemplary method for selecting a DRC profile. [Figure 6b] This is the second half of a flowchart illustrating an exemplary method for selecting a DRC profile. [Modes for carrying out the invention]

[0015] As described above, this paper addresses the technical challenge of enabling audio content designers and / or distributors to control the quality and intelligibility of audio content for various types of rendering modes. An exemplary rendering mode is the home theater rendering mode, where audio content is played using transducers that typically allow for a very wide dynamic range in a quiet environment. Another exemplary rendering mode is the flat panel mode, where audio content is played using transducers, such as those in a TV set, which typically allow for a reduced dynamic range compared to a home theater. A further exemplary rendering mode is the portable speaker mode, where audio content is played using speakers from a portable electronic device (such as a smartphone). The dynamic range in this rendering mode is typically smaller than that of the rendering modes described above, and the environment is often noisy. Another exemplary rendering mode is the portable headphone mode, where audio content is played using headphones associated with a portable electronic device. The dynamic range is limited, but typically higher than that provided by the speakers of a portable electronic device.

[0016] To allow for high quality and high intelligibility for various rendering modes, various DRC (Dynamic Range Control) profiles for various rendering modes may be provided along with the audio content. The audio content may be transmitted in a sequence of frames. The sequence of frames may include I (i.e., independent) frames that can be decoded independently of preceding or subsequent frames. Further, the sequence of frames may typically include other types of frames (e.g., P and / or B frames) that exhibit dependencies on preceding and / or subsequent frames. At least some of the frames in the sequence of frames may include multiple different DRC profiles for multiple different rendering modes. In particular, the I frames of the sequence of frames may include the plurality of DRC profiles.

[0017] By inserting multiple different DRC profiles into the sequence of audio frames, the audio decoder can select an appropriate DRC profile for a particular rendering mode. As a result, it can be guaranteed that the rendered audio signal has high quality (especially without clipping or distortion introduced by the transducer) and high intelligibility.

[0018] In the following, various aspects of dynamic range control are described. Without customized dynamic range control, input audio information (e.g., PCM samples, time-frequency samples in a QMF matrix, etc.) is often reproduced at a loudness level that is not appropriate for the particular playback environment of the playback device (i.e., including the physical and / or mechanical playback limits of that device). This is because the particular playback environment of the playback device may be different from the target playback environment in which the encoded audio content was encoded at the encoding device.

[0019] The techniques described in this document can be used to support dynamic range control of a wide variety of audio content customized for any of a wide variety of playback environments, while maintaining the perceptual quality of the audio content and while maintaining the artist's intent to adapt the content to various playback environments.

[0020] Dynamic range control (DRC) refers to a time-varying, level-dependent audio processing operation that changes (e.g., compresses, cuts, expands, boosts, etc.) a signal to convert the input dynamic range of loudness levels in audio content to an output dynamic range different from the input dynamic range. For example, in a certain dynamic range control scenario, a small sound may be mapped (e.g., boosted, etc.) to a higher loudness level, and a large sound may be mapped (e.g., cut, etc.) to a lower loudness value. As a result, in the loudness domain, the output range of loudness levels is, in this example, smaller than the input range of loudness levels. However, in some embodiments, dynamic range control may be reversible such that the original range can be restored. For example, as long as the mapped loudness level in the output dynamic range mapped from the original loudness level is below the clipping level and each unique original loudness level is mapped to a unique output loudness level, an expansion operation can be performed to restore the original range.

[0021] The DRC techniques described in this paper can be used to provide a better listening experience in certain playback environments or situations. For example, quiet sounds in a noisy environment may be masked by noise that makes them inaudible. Conversely, loud sounds may be undesirable in certain situations, such as when disturbing neighbors (e.g., within a "late-night" listening mode). Many devices, typically with speakers of small form factor, cannot reproduce sound at high output levels or without perceptible distortion. In some cases, lower signal levels may be reproduced below the human auditory threshold. DRC techniques can perform mapping of input loudness levels to output loudness levels based on DRC gains (scaling factors such as scaling audio amplitude, boosting ratios, or cutting ratios) found using dynamic range compression curves.

[0022] A dynamic range compression curve is a function (e.g., a lookup table, curve, or multi-segment piecewise linear function) that maps individual input loudness levels (e.g., sounds other than dialogue) determined from individual audio data frames to corresponding output loudness levels, and consequently, to individual gains (one or more) for dynamic range control to convert those input loudness levels to the corresponding output loudness levels. Each individual gain indicates the amount of gain that should be applied to the signal to map the corresponding individual input loudness level to the intended output loudness level. The output loudness level after applying the individual gains represents the target loudness level for the audio content in individual audio data frames in a particular playback environment.

[0023] In addition to specifying a mapping between gain and loudness level, the dynamic range compression curve may include, or may provide, specific release and attack times when applying a particular gain. Attack refers to the increase in signal energy (or loudness) between consecutive time samples, while release refers to the decrease in energy (or loudness) between consecutive time samples. The attack time (e.g., 10 milliseconds, 20 milliseconds, etc.) is the time constant used to smooth the DRC gain when the corresponding signal is in attack mode. The release time (e.g., 80 milliseconds, 100 milliseconds, etc.) is the time constant used to smooth the DRC gain when the corresponding signal is in release mode. In some embodiments, additionally, optionally, or alternatively, these time constants are used to smooth the signal energy (or loudness) prior to determining the DRC gain.

[0024] Different dynamic range compression curves may correspond to different playback environments (i.e., different rendering modes). For example, a dynamic range compression curve for a flat-panel TV playback environment may differ from a dynamic range compression curve for a portable device playback environment. For example, a first dynamic range compression curve for a first playback environment of a portable device with speakers may differ from a second dynamic range compression curve for a second playback environment of the same portable device with a headset.

[0025] Figure 1 shows a block diagram of exemplary components of the audio decoder 100. The audio decoder 100 includes a data extractor 104, a dynamic range controller 106, and an audio renderer 108. The data extractor 104 is configured to receive an encoded input signal 102. The encoded input signal 102 described in this paper may be a bitstream containing encoded (e.g., compressed) input audio data frames (particularly sequences of audio frames) and possibly metadata. The bitstream may be an AC-4 bitstream. The data extractor 104 is configured to extract / decode input audio data frames and metadata from the encoded input signal 102. Each input audio data frame has multiple encoded audio data blocks, each representing multiple audio samples. Each frame represents a (e.g., constant) time interval containing a certain number of audio samples. The frame size may vary with the sample rate and the encoded data rate. An audio sample is a quantized audio data element (e.g., an input PCM sample, an input time-frequency sample in a QMF matrix) that represents spectral content in one, two, or more (audio) frequency bands or frequency ranges. A quantized audio data element within an input audio data frame may represent a sound pressure wave in the digital (quantized) domain. A quantized audio data element may cover a finite range of loudness levels below a possible maximum value (e.g., a clipping level, maximum loudness level).

[0026] Metadata can be used by the audio decoder 100 to process the input audio data frame. The metadata may include various operational parameters relating to one or more operations to be performed by the decoder 100, one or more dynamic range compression curves (i.e., one or more DRC profiles), normalization parameters relating to the dialogue loudness level represented in the input audio data frame, and so on. Dialogue loudness level can refer to levels such as dialogue loudness, program loudness, average dialogue loudness, etc. (e.g., psychoacoustic, perceptual, etc.) in the entire program (e.g., a film, television program, radio broadcast, etc.), a part of the program, or the dialogue of the program.

[0027] Some or all of the operation and functions of the decoder 100 or its modules (e.g., data extractor 104, dynamic range controller 106, etc.) may be adapted in response to metadata extracted from the encoded input signal 102. For example, metadata—including but not limited to dynamic range compression curves, dialogue loudness levels, etc.—may be used by the decoder 100 to generate digital domain output audio data elements (e.g., output PCM samples, output time-frequency samples in a QMF matrix, etc.). The output data elements can then be used to drive audio channels or speakers to achieve a specified loudness or reference playback level during playback in a particular playback environment.

[0028] The dynamic range controller 106 may be configured to receive some or all of the audio data elements and metadata in the input audio data frame and to perform audio processing operations (e.g., dynamic range control operations, gain smoothing operations, gain limiting operations, etc.) on the audio data elements in the input audio data frame based on metadata extracted from the encoded audio signal 102, at least partially.

[0029] In particular, the dynamic range controller 106 may include a selector 110, a loudness calculator 112, and a DRC gain unit 114. The selector 110 may be configured to determine a speaker configuration related to a specific playback environment in the decoder 100 (e.g., home theater mode, flat panel mode, portable device mode with speakers, portable device mode with headphones, 5.1 speaker configuration mode, 7.1 speaker configuration mode, etc.). Furthermore, the selector 110 may be configured to select a specific dynamic range compression curve (i.e., a certain DRC profile) from various dynamic range compression curves (i.e., from the plurality of DRC profiles) extracted from the metadata of the encoded input signal 102.

[0030] The loudness calculator 112 may be configured to calculate one or more types of loudness levels represented by the audio data elements in the input audio data frame. Examples of loudness level types, but not limited to, include, but any, individual loudness levels across individual frequency bands in individual channels over individual time intervals, broadband loudness levels over a wide frequency range in individual channels, loudness levels determined from or smoothed over an audio data block or frame, loudness levels determined from or smoothed over two or more audio data blocks or frames, and loudness levels smoothed over one or more time intervals. Zero, one or more of these loudness levels may be modified by the decoder 100 for dynamic range control.

[0031] To determine the loudness level, the loudness calculator 112 can determine one or more time-dependent physical sound wave attributes, such as spatial and / or local pressure levels at a particular audio frequency, which are represented by the audio data elements in the input audio data frame. The loudness calculator 112 can use the one or more time-varying physical wave attributes to derive one or more types of loudness levels based on one or more psychoacoustic functions that model human loudness perception. The psychoacoustic function may be a nonlinear function—constructed based on a model of the human auditory system—that translates / maps a particular spatial pressure level at a particular audio frequency to a specific loudness for that particular audio frequency.

[0032] Loudness levels (e.g., broadband, wideband) across multiple (audio) frequencies or frequency bands may be derived through the integration of specific loudness levels across multiple (audio) frequencies or frequency bands. Time-averaged, smoothed, etc., loudness levels over one or more time intervals (e.g., longer than represented by audio data blocks or audio data elements in frames) may be obtained using one or more smoothing filters implemented as part of the audio processing operation in decoder 100. Another exemplary method for determining the (broadband) loudness level is specified in ITU-R BS.1770. The method specified in ITU-R BS.1770 applies time-domain filtering to a time-domain input audio signal, then calculates the RMS (root mean square) level for each channel of the input audio signal, and then integrates and gates the resulting loudness levels across the channels.

[0033] Specific loudness levels for different frequency bands may be calculated for each audio data block (e.g., 256 samples). Pre-filters may be used to apply frequency weighting (e.g., similar to IEC B weighting) to the specific loudness levels in integrating them to a broadband loudness level. The sum of broad loudness levels across two or more channels (e.g., front left, front right, center, left surround, right surround) may be performed to provide an overall loudness level for those two or more channels.

[0034] The overall loudness level may refer to the broadband loudness level of a single channel (e.g., the center) in a speaker configuration. The overall loudness level may refer to the broadband loudness level of multiple channels. The multiple channels may be all channels in a speaker configuration (i.e., for a given rendering mode). Additionally, optionally, or alternatively, the multiple channels may include subsets of channels in a speaker configuration (e.g., subsets of channels including left front, right front, and low-frequency effects (LFE); subsets of channels including left surround and right surround; subsets of channels including the center, etc.).

[0035] Loudness levels (e.g., broadband, wideband, overall, specific) may be used as input to find the corresponding DRC gain (e.g., static, unsmoothed, unlimited) from a selected dynamic range compression curve. The loudness levels used as input to find the DRC gain may first be adjusted or normalized with respect to the dialogue loudness level from metadata extracted from the encoded audio signal 102 and / or the output reference level in rendering mode. Adjustments and normalization with respect to the dialogue loudness level / output reference level may be performed on the portion of the audio signal in the encoded audio signal 102 in the non-loudness region (e.g., the SPL region) before the specific spatial pressure level represented in the portion of the audio content in the encoded audio signal 102 is converted to the specific loudness level of that portion of the audio content in the encoded audio signal 102.

[0036] The DRC gain unit 114 may be configured with a DRC algorithm to generate gain (for example, for dynamic range control, gain limiting, gain smoothing, etc.) and apply the gain to one or more loudness levels in one or more types of loudness levels represented by audio data elements in the input audio data frame to achieve a target loudness level for that particular playback environment. The application of gain (e.g., DRC gain) as described in this paper may occur in the loudness domain. For example, the gain may be generated based on a loudness calculation (which may be expressed in SPL compensated for a thorn or simply for example, an unconverted dialogue loudness level), smoothed, and applied directly to the input signal. Techniques as described in this paper may calculate the corresponding gain to be applied to the signal by applying the gain to the signal in the loudness domain, then converting the signal from the loudness domain to the original (linear) SPL domain, and evaluating the signal before and after the gain was applied to the signal in the loudness domain. Then, the ratio (or difference, if expressed in logarithmic dB terms) determines the corresponding gain for that signal.

[0037] The DRC algorithm may operate with multiple DRC parameters. These DRC parameters include a dialogue loudness level, which is already calculated by an upstream encoder 150 (described, for example, in the context of Figure 2) and embedded in the encoded audio signal 102, and can be retrieved by the decoder 100 from the metadata in the encoded audio signal 102. The dialogue loudness level from the upstream encoder 150 represents the average dialogue loudness level (e.g., per program, relative to the energy of a full-scale 1kHz sine wave, relative to the energy of a reference square wave, etc.). The dialogue loudness level extracted from the encoded audio signal 102 may be used to reduce differences in loudness levels between programs. The reference dialogue loudness level may be set to the same value between different programs in the same specific playback environment at the decoder 100. Based on the dialogue loudness level from metadata, the DRC gain unit 114 may apply a dialogue loudness-related gain to each audio data block in the program so that the average output dialogue loudness level (or output reference level) across multiple audio data blocks of the program is raised / lowered to a reference dialogue loudness level for that program (e.g., a pre-configured, system-default, user-configurable, or profile-dependent level). The dialogue loudness level may be used to calibrate the DRC algorithm. In particular, the null band of the DRC algorithm may be adjusted to the dialogue loudness level. Alternatively, the desired output reference level may be used to calibrate the DRC algorithm when the DRC algorithm is applied to a signal to which gain has been applied to change the dialogue loudness level to be equal to the desired output reference level. The dialogue loudness level may correspond to a so-called dialnorm parameter, if speech gating has been applied to determine the dialnorm parameter.In some embodiments, the dialogue loudness level corresponds to a `dialnorm` parameter determined by gating based on a loudness level threshold, rather than by using speech gating.

[0038] DRC gain may be used to address differences in loudness levels within a program by boosting or cutting signal portions in soft and / or loud sounds according to a selected dynamic range compression curve. One or more of these DRC gains may be calculated / determined by a DRC algorithm based on a selected dynamic range compression curve determined from one or more corresponding audio data blocks, audio data frames, etc., and loudness levels (e.g., broadband, wideband, overall, specific, etc.).

[0039] The loudness level used to determine the DRC gain (e.g., static, unsmoothed, ungain-limited, etc.) by searching for a selected dynamic range compression curve may be calculated over a short interval (e.g., about 5.3 milliseconds). The integration time of the human auditory system (e.g., about 200 milliseconds) can be much longer. The DRC gain obtained from the selected dynamic range compression curve may be smoothed with a time constant to account for the long integration time of the human auditory system. To implement a fast rate of change (increase or decrease) in the loudness level, a short time constant may be used to cause the change in loudness level over a short time interval corresponding to a short time constant. Conversely, to implement a slow rate of change (increase or decrease) in the loudness level, a long time constant may be used to change the loudness level over a long time interval corresponding to a long time constant.

[0040] The human auditory system may respond to increasing and decreasing loudness levels with different integration times. To smooth the static DRC gain retrieved from a selected dynamic range compression curve, different time constants may be used depending on whether the loudness level is increasing or decreasing. For example, corresponding to the characteristics of the human auditory system, attack (increase in loudness level) may be smoothed with a relatively short time constant (e.g., attack time), while release (decrease in loudness level) may be smoothed with a relatively long time constant (e.g., release time).

[0041] The DRC gain for a portion of the audio content (e.g., one or more audio data blocks, audio data frames, etc.) may be calculated using the loudness level determined from the aforementioned portion of the audio content. The loudness level to be used for searching in the selected dynamic range compression curve may first be adjusted with respect to (or in relation to) the dialogue loudness level (e.g., of a program, etc., in which the audio content is part) in the metadata extracted from the encoded audio signal 102.

[0042] Reference dialog loudness level / output reference level (for example, -31dB in "line" mode) FS In "RF" mode, -20dB FS (etc.) may be specified or established for a specific playback environment in decoder 100. Additionally, alternatively, or optionally, in some embodiments, the user may be given control over setting or changing the reference dialog loudness level in decoder 100.

[0043] The DRC gain unit 114 may be configured to determine the dialogue loudness relation gain for audio content, causing a change from the input dialogue loudness level to a reference dialogue loudness level as the output dialogue loudness level.

[0044] The audio renderer 108 may be configured to apply a gain determined based on DRC, gain limiting, gain smoothing, etc., to the input audio data extracted from the encoded audio signal 102, and then generate channel-specific audio data 116 (e.g., multi-channel) for that particular speaker configuration. The channel-specific audio data 116 may be used to drive speakers, headphones, etc., represented in that speaker configuration.

[0045] Additionally and / or optionally, the decoder 100 may be configured to perform one or more other operations related to processing, rendering, downmixing, resampling, etc., of the input audio data.

[0046] The techniques described in this paper can be used with a variety of speaker configurations to accommodate diverse surround sound configurations (e.g., 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10-60 speaker configurations, 60+ speaker configurations, object signals or combinations of object signals, etc.) and a variety of different rendering environment configurations (e.g., cinemas, parks, opera houses, concert halls, bars, homes, auditoriums, etc.).

[0047] Figure 2 shows an exemplary encoder 150. The encoder 150 may include an audio content interface 152, a dialogue loudness analyzer 154, a DRC reference storage unit 156, and an audio signal encoder 158. The encoder 150 may be part of a broadcasting system, an internet-based content server, an over-the-air network operator system, a film production system, etc.

[0048] The audio content interface 152 may be configured to receive audio content 160 and audio content control input 162, and to generate an encoded audio signal 102 based at least partially on some or all of the audio content 160 and audio content control input 162. For example, the audio content interface 152 may be used to receive audio content 160 and audio content control input 162 from a content creator, content provider, etc.

[0049] Audio content 160 may consist of audio only, or part or all of overall media data including audiovisuals, etc. Audio content 160 may include one or more of the following: parts of a program, a program, several programs, one or more commercials, etc.

[0050] The dialogue loudness analyzer 154 may be configured to determine / establish one or more dialogue loudness levels for one or more parts of the audio content 152 (e.g., one or more programs, one or more commercials, etc.). The audio content may be represented by one or more sets of audio tracks. The dialogue audio content of the audio content may be on a separate audio track, and / or at least a portion of the dialogue audio content of the audio content may be on an audio track that also contains non-dialogue audio content.

[0051] The audio content control input 162 may include some or all of the following: user control inputs, control inputs provided by external systems / devices to the encoder 150, control inputs from content creators, and control inputs from content providers. For example, a user such as a mixing engineer may provide / specify one or more dynamic range compression curve identifiers. These identifiers may be used to retrieve one or more dynamic range compression curves that best fit the audio content 160 from a data storage unit such as a DRC reference storage unit (156).

[0052] The DRC reference storage unit 156 may be configured to store DRC reference parameter sets, etc. These DRC reference parameter sets may include definition data for one or more dynamic range compression curves, etc. The encoder 150 may encode two or more dynamic range compression curves into the encoded audio signal 102 (for example, concurrently). Zero, one or more of these dynamic range compression curves may be standard-based, proprietary, customized, or modifiable by the decoder. For example, the dynamic range compression curves in Figures 3 and 4 may be encoded into the encoded audio signal 102 (for example, concurrently).

[0053] The audio signal encoder 158 may be configured to receive audio content from the audio content interface 152 and dialogue loudness levels from the dialogue loudness analyzer 154, retrieve one or more DRC reference parameter sets (i.e., DRC profiles) from the DRC reference storage unit 156, format the audio content into audio data blocks / frames, format the dialogue loudness levels, DRC reference parameter sets, etc., into metadata (e.g., metadata containers, metadata fields, metadata structures, etc.), and encode the audio data blocks / frames and metadata into an encoded audio signal 102. The audio content to be encoded in the encoded audio signal as described in this paper may be received in one or more of a variety of source audio formats by one or more of a variety of methods, such as wirelessly, via a wired connection, via a file, or via an internet download.

[0054] The encoded audio signal 102 described in this paper may be part of an overall media data bitstream (for example, for audio broadcasts, audio programs, audiovisual programs, audiovisual broadcasts, etc.). The media data bitstream may be accessed from servers, computers, media storage devices, media databases, media files, etc. The media data bitstream may be broadcast, transmitted, or received over one or more wireless or wired network links. The media data bitstream may be communicated over one or more intermediaries such as network connections, USB connections, wide area networks, local area networks, wireless connections, optical connections, buses, crossbar connections, serial connections, etc.

[0055] Any of the components depicted (for example in Figures 1 and 2) may be implemented as one or more processes and / or one or more IC circuits (e.g., ASICs, FPGAs, etc.) in hardware, software, or a combination of hardware and software.

[0056] Figures 3 and 4 show exemplary dynamic range compression curves that can be used by the DRC gain unit 104 in decoder 100 to derive DRC gain from the input loudness level. As shown in the figures, the dynamic range compression curve may be centered on a reference loudness level in the program (e.g., output reference level) to provide an appropriate overall gain for a particular playback environment. Exemplary definition data for the dynamic range compression curve (e.g., in the metadata of the encoded audio signal 102) (e.g., including, but not limited to, boost ratio, cut ratio, attack time, release time, etc.) is shown in the table below. Different profiles (e.g., film standard, film light, music standard, music light, speech, etc.) represent different playback environments (e.g., in decoder 100).

[0057] [Table 1] dB SPL or dB FS Loudness level and dB expressed as SPL One or more compression curves described using gains expressed in dB with respect to may be received. On the other hand, DRC gain calculation is dB SPL The process is performed using a different loudness representation (e.g., Thorne) that has a nonlinear relationship with the loudness level. In this case, the compression curve used in the DRC gain calculation may be transformed so that it is described using the different loudness representation (e.g., Thorne).

[0058] Figure 5 shows an exemplary encoded audio signal 102 containing a sequence of frames (numbered from n+1 to n+30, where n is an integer). In the illustrated example, every fifth frame is an I-frame. In the illustrated example, I-frames (n+1) have multiple DRC profiles (identified as AVR (Audio / Video Receiver) for Home Theater, Flat Panel, Portable HP (Headphones), and Portable SP (Speakers)). Each DRC profile has a dynamic range compression curve as shown in Figures 3 and 4.

[0059] The aforementioned DRC profiles can be repeatedly inserted within the I-frame of a frame sequence. This allows the decoder 100 to determine the appropriate DRC profile for the encoded audio signal 102 and the current rendering mode at startup of the decoded audio signal 102, when tuning into the broadcast audio program, and / or after the junction. On the other hand, the iterative transmission of the entire set of DRC profiles leads to relatively high bitstream overhead. In view of this, it is proposed to transmit a changing subset of the DRC profiles within the I-frame of the encoded audio signal 102.

[0060] Figure 5 shows an example of inserting a DRC profile within a sequence of frames. In the illustrated example, only a single DRC profile from a complete set of DRC profiles is inserted into the I-frame. The DRC profile inserted into the I-frame changes with each I-frame, and as a result, after N I-frames (N=4 in the illustrated example), the decoder 100 will have received a complete set of N DRC profiles. This reduces the data rate required to transmit a complete set of DRC profiles while ensuring that the decoder 100 receives a complete set of DRC profiles within a reasonable time.

[0061] Figures 6a and 6b show flowcharts of an exemplary method 600 for determining a DRC profile for decoding a frame of an encoded audio signal 102. Method 600 may be performed by a decoder 100 (particularly a selector 110). Upon commencement of receiving the encoded audio signal 102, the DRC profile to be used by the decoder 100 may be initialized. The DRC profile used to decode the current frame of the encoded audio signal 102 may be referred to as the current DRC profile. Thus, at startup, the current DRC profile may be initialized. In particular, the default DRC profile (available in the decoder 100) may be set to be the current DRC profile used to render the current frame (step 601 of the method). Thus, the variable "profile" may be set to the default DRC profile (profile=default_DRC_profile). Furthermore, the decoder 100 may track previously used profiles. Previously used profiles may be set to undefined (prev_profile=undefined).

[0062] Method 600 may further include step 602, which takes a new frame (i.e., the current frame) to be decoded from the encoded audio signal 102. In step 603, it is verified whether the new frame is an I-frame that may contain a DRC profile. If the new frame is not an I-frame, Method 600 proceeds to step 604, which processes the new frame using the current DRC profile. Furthermore, in step 605, the previously used profile is set as the current DRC profile (prev_profile=profile).

[0063] If the new frame is an I-frame, step 606 of the method may check whether the I-frame contains DRC data. For example, the metadata of the I-frame may include a flag indicating whether the I-frame contains DRC data. If no DRC data is present, the method may proceed to steps 604 and 605. Otherwise, the method may proceed to step 607.

[0064] In step 607 of the method, it may be verified whether the new frame is the first frame of the encoded audio signal 102 to be decoded. As can be seen from the flowcharts in Figures 6a and 6b, this can be verified by checking the prev_profile variable. If the prev_profile variable is undefined, the new frame is the first frame to be decoded. If the new frame is the first frame to be decoded, the decoder 100 may use a predefined DRC profile other than the default DRC profile. For this purpose, the metadata of the new frame may include an identifier (ID) for such a predefined DRC profile. Such predefined DRC profiles may be stored in a database in the decoder 100. The use of predefined DRC profiles can provide a bitrate-efficient means of signaling the DRC profile to be used to the decoder 100, because only the ID of the predefined profile needs to be transmitted (step 608 of the method). Predefined DRC profiles that are signaled using their IDs may be called implicit DRC profiles.

[0065] In some cases, it may be beneficial to use only a single predefined DRC profile other than the default DRC profile. In such cases, the decoder 100 may be configured to set the profile variable to that predefined (i.e., implicit) DRC profile without receiving any IDs in the metadata of the new frame.

[0066] Method 600 may further include verifying whether the metadata of a new frame contains one or more explicit DRC profiles (step 609). An explicit DRC profile includes an ID for identifying the explicit DRC profile. Furthermore, an explicit DRC profile typically includes definitional data for a dynamic range compression curve, such as those shown in Figures 3 and 4. The dynamic range compression curve may be defined as a piecewise linear function. Furthermore, an explicit DRC profile may indicate a range of output reference levels (ORL) to which the explicit DRC profile is applicable. For example, a default DRC profile and / or the predefined (implicit) DRC profiles may be applicable to output reference levels ranging from -31 dB FS to 0 dB FS.

[0067] The ORL of a rendering device may indicate its dynamic range capability. Typically, the dynamic range capability decreases as the ORL increases. For high ORL, a compression curve with a high degree of compression should be used to render the audio signal in an intelligible manner without clipping. On the other hand, for low ORL, compression may be reduced to render the audio signal with a high dynamic range. Due to the high dynamic range capability of the rendering device, the intelligibility of the audio signal is still guaranteed.

[0068] If the metadata for a new frame contains at least one explicit DRC profile, the profile data of the first DRC profile is read (step 610). Furthermore, it is verified whether the ORL range of the first DRC profile is applicable to the rendering device currently in use (step 611). If not, method 600 proceeds to search for another explicit DRC profile within the metadata for the new frame. On the other hand, if an explicit DRC profile is applicable to the rendering device, this explicit DRC profile may be set as the current DRC profile to be used to process the new frame (step 614).

[0069] Method 600 may further include verifying whether a headphone rendering mode is being used and whether an explicit DRC profile is applicable to the headphone rendering mode (step 612). Furthermore, Method 600 may also include verifying whether the explicit DRC profile is an updated profile compared to a previously used profile (step 613). For this purpose, the ID of the explicit DRC profile may be compared with the ID of the currently used profile. This ensures that the decoder 100 always uses the most recent current DRC profile.

[0070] Using method 600, it can be ensured that the decoder 100 will always identify a DRC profile for rendering frames of the encoded audio signal 102, even if the data has not yet received a DRC profile for the current rendering mode (i.e., for the current rendering device). Furthermore, it can be ensured that the DRC profile for the current rendering mode is applied as soon as the decoder 100 receives the corresponding DRC profile.

[0071] Therefore, a method 600 for decoding the encoded audio signal 102 is described. The encoded audio signal 102 has a sequence of frames. Furthermore, the encoded audio signal 102 exhibits multiple different Dynamic Range Control (DRC) profiles for multiple corresponding different rendering modes. Examples for different rendering modes (or different playback environments) are a first DRC profile for use in home theater rendering mode; a second DRC profile for use in flat panel rendering mode; a third DRC profile for use in portable device speaker rendering mode; and / or a fourth DRC profile for use in headphone rendering mode. The DRC profile defines a particular DRC behavior. The DRC behavior may be described by a compression curve (and time constants) and / or DRC gain. The DRC gains may be temporally equidistant gains that can be applied to the encoded audio signal 102 to deploy the DRC. The compression curve may be accompanied by time constants that together constitute a DRC algorithm. DRC typically reduces the volume of loud sounds and amplifies quiet sounds, thereby compressing the dynamic range of the audio signal for an improved experience in less-than-ideal playback environments.

[0072] A sequence of frames typically includes multiple consecutive frames that constitute an audio signal. An audio program (for example, a television or radio program to be broadcast) may include multiple audio signals that are linked together at junctions. For example, the main audio program may be repeatedly interrupted by commercial breaks. A sequence of frames may correspond to a complete audio program. Alternatively, a sequence of frames may correspond to one of the multiple audio signals that constitute a complete audio program.

[0073] Different subsets of DRC profiles from the aforementioned multiple DRC profiles may be included within different frames of a sequence of frames. This allows two or more frames of the sequence of frames to consolidate and include the aforementioned multiple DRC profiles. As described above, the distribution of DRC profiles across multiple frames of a sequence of frames leads to a reduction in the bitstream overhead for signaling the aforementioned multiple DRC profiles.

[0074] Method 600 may include determining a first rendering mode from the plurality of different rendering modes. In particular, it may be determined which rendering mode is used to render the encoded audio signal 102. Furthermore, Method 600 may include determining one or more DRC profiles from the plurality of DRC profiles contained in the current frame of the sequence of frames 609, 610. In other words, one or more DRC profiles may be determined from a subset of DRC profiles contained in the current frame. Furthermore, it may be determined whether at least one of the one or more DRC profiles is applicable to the first rendering mode 611. Determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode 611 may include determining a first output reference level for the first rendering mode, determining the range of output reference levels to which the DRC profiles from the one or more DRC profiles are applicable, and determining whether the first output reference level falls within that range of output reference levels.

[0075] Method 600 may further include selecting a default DRC profile as the current DRC profile if none of the one or more DRC profiles are applicable to the first rendering mode. The definition data of the default DRC profile is typically known in the decoder for decoding the encoded audio signal 102. Furthermore, Method 600 may include decoding (and / or rendering) the current frame using the current DRC profile. Thus, even if the decoder 100 has not yet received a DRC profile specific to the encoded audio signal 102, it can be ensured that the decoder 100 will utilize a DRC profile (and dynamic range compression curve).

[0076] Alternatively or additionally, method 600 may include selecting a first DRC profile from the one or more DRC profiles as the current DRC profile if it is determined that the first DRC profile is applicable to the first rendering mode 604. As a result, the decoder 100 is configured to use the first DRC profile that is best suited for the encoded audio signal 102 and the first rendering mode as soon as the decoder 100 receives the first DRC profile.

[0077] Method 600 may further include determining whether the current frame of a sequence of frames contains one or more DRC profiles from the plurality of DRC profiles, i.e., whether the current frame contains a subset of DRC profiles 603, 606. As outlined in the context of Figure 5, a subset of DRC profiles is typically contained within an I-frame of a sequence of frames. Thus, determining whether the current frame contains one or more DRC profiles from the plurality of DRC profiles, or whether the current frame contains a subset of DRC profiles 603, 606 may include determining whether the current frame is an I-frame 603. As described above, an I-frame may be a frame that can be decoded independently of any other frame from a sequence of frames. This may be due to the fact that the data contained in such an I-frame is transmitted in a manner independent of data from previous or subsequent frames. In particular, the data in an I-frame is not differentially encoded with respect to data contained in previous or subsequent frames.

[0078] Furthermore, determining whether the current frame contains one or more DRC profiles from the plurality of DRC profiles, or whether the current frame contains a subset of DRC profiles, 603, 606 may include verifying the DRC profile flags contained within the current frame, 606. The DRC profile flags in the bitstream of the encoded audio signal provide a bandwidth-efficient and computationally efficient means for identifying frames that carry DRC profiles.

[0079] Method 600 may further include determining whether the current frame represents a particular implicit DRC profile from among several implicit DRC profiles. The implicit DRC profile may include a predefined legacy compression curve and time constant that can be used to transcode to E-AC-3. As described above, the definition data of the implicit DRC profile may be known in the decoder 100 for decoding the input audio signal 102. In contrast to the default DRC profile, the implicit DRC profile may be specific to various types of audio signals (e.g., as shown in Table 1). The current frame in a sequence of frames may represent a particular implicit DRC profile (e.g., using an identifier ID). This may provide a bandwidth-efficient means of signaling an appropriate DRC profile for the encoded audio signal 102. The implicit DRC profile may be selected as the current DRC profile if it is determined that the current frame represents an implicit DRC profile 608.

[0080] Decoding the current frame may include leveling the sequence of frames to the first output reference level of the first rendering mode. Furthermore, decoding the current frame may include adapting the loudness level of the current frame using the dynamic range compression curve specified in the current DRC profile. The adaptation of the loudness level may be performed as outlined in the context of Figure 1.

[0081] Depending on the number of frames from the frame sequence, the current DRC profile may correspond to a default DRC profile (which is typically independent of the input audio signal 102), an implicit DRC profile (which may be adapted to the input audio signal 102 in a limited way), or the first explicit DRC profile (which may be designed for the input audio signal 102 and / or the first rendering mode).

[0082] Typically, only a subset of frames contains a DRC profile. Once a current DRC profile is selected, it may be maintained to decode frames in the sequence that do not contain any DRC profiles. Furthermore, even if frames with a DRC profile are received, the current DRC profile may be maintained unless a DRC profile that is more relevant to a newer and / or encoded audio signal 102 than the current DRC profile is received (where the selected first explicit DRC profile is more relevant than the selected implicit DRC profile, and the selected implicit DRC profile is more relevant than the default DRC profile). This ensures the continuity and optimality of the DRC profiles used.

[0083] As a complement to method 600 for decoding an encoded audio signal 102, a method for generating or encoding an encoded audio signal 102 is described. The encoded audio signal 102 has a sequence of frames. Furthermore, the encoded audio signal 102 represents multiple different dynamic range control (DRC) profiles for multiple different rendering modes. The method includes inserting different subsets of DRC profiles from the multiple DRC profiles into different frames of the sequence of frames so that two or more frames of the sequence of frames congruently contain the multiple DRC profiles. In other words, fewer subsets of DRC profiles than the total number of DRC profiles may be provided with the various frames of the sequence of frames. In this way, the overhead of the encoded audio signal 102 can be reduced while providing the corresponding decoder 100 with a complete set of DRC profiles. In other words, the advantage of this technique is that the encoder 150 has increased degrees of freedom in how it transmits the DRC data. This degree of freedom can be used to reduce the bitrate.

[0084] A sequence of frames may include a subsequence of I-frames (for example, every Xth frame in the sequence of frames may be an I-frame). Various subsets of the DRC profile may be inserted into various (e.g., successive) I-frames of the I-frame subsequence. To further reduce bandwidth, I-frames may be skipped; that is, some I-frames may not contain DRC profile data.

[0085] A subset of DRC profiles (for example, each subset) may contain only a single DRC profile. In particular, the plurality of DRC profiles may contain N DRC profiles, where N is an integer and N > 1. The N DRC profiles may be inserted into N different frames from the sequence of frames. This can minimize the bitrate required for transmitting the DRC profiles.

[0086] The method may further include inserting all of the plurality of DRC profiles into a first frame of the sequence of frames (for example, in the first frame of the sequence of frames of an audio signal). As a result, rendering of the encoded audio signal 102 can be started directly with the correct explicit DRC profile. As described above, the audio program may be divided into multiple sub-audio programs. For example, the main audio program may be interrupted by a commercial pause. It may be beneficial to insert all of the plurality of DRC profiles into the first frame of each sub-audio program. In other words, it may be beneficial to insert all of the plurality of DRC profiles immediately after one or more junctions in an audio program containing multiple sub-audio programs.

[0087] Various subsets of DRC profiles from the aforementioned plurality of DRC profiles may be inserted into different frames of the sequence of frames. This allows M directly consecutive subsets from the sequence of frames to congruently contain the plurality of DRC profiles, where M is an integer and M > 1. In other words, the plurality of DRC profiles may be transmitted repeatedly within a block of M frames. As a result, the decoder 100 may need to wait up to M frames to obtain the optimal explicit DRC profile for the encoded audio signal 102.

[0088] The method may further include inserting a flag into the frame of the sequence of frames, where the flag indicates whether or not the frame contains a DRC profile. By providing such a flag, the corresponding decoder 100 can efficiently identify frames containing DRC profile data.

[0089] The DRC profiles of the aforementioned plurality of DRC profiles may be explicit DRC profiles that include (i.e., carry) definition data for defining the dynamic range compression curve.

[0090] As outlined in this paper, the dynamic range compression curve provides a mapping between input loudness and output loudness and / or gain to be applied to the audio signal. In particular, the definition data may include one or more of the following: a boost gain for boosting input loudness; a boost gain range indicating the range of input loudness to which the boost gain is applicable; a null band range indicating the range of input loudness to which a gain of 0 dB is applicable; a cut gain for attenuating input loudness; a cut gain range indicating the range of input loudness to which the cut gain is applicable; a boost gain ratio indicating the transition between the null gain and the boost gain; and / or a cut gain ratio indicating the transition between the null gain and the cut gain.

[0091] The method may further include inserting an implicit DRC profile instruction (e.g., identifier, ID), where the implicit DRC profile definition data is typically known to the decoder 100 of the encoded audio signal 102. The implicit DRC profile instruction can provide a bandwidth-efficient means for signaling a DRC profile that is (in a limited way) adapted to the encoded audio signal 102.

[0092] As outlined above, the frames in the aforementioned sequence of frames typically contain audio data and metadata. A subset of the DRC profile is typically inserted as metadata.

[0093] A DRC profile may include definitional data that defines the range of output reference levels to which the DRC profile is applicable. Typically, the output reference level indicates the dynamic range of a given rendering mode. In particular, the dynamic range of a rendering mode may decrease with increasing output reference levels, and vice versa. Furthermore, the maximum boost and maximum cut gains of the dynamic range compression curve of the DRC profile may increase with increasing output reference levels, and vice versa. Therefore, the output reference level provides an efficient means of selecting an appropriate DRC profile (with an appropriate dynamic range compression curve) for a particular rendering mode.

[0094] The method may further include generating a bitstream containing the encoded audio signal 102. The bitstream may be an AC4 bitstream; that is, the bitstream may conform to the AC4 bitstream format.

[0095] The method may further include inserting explicit DRC gains for the encoded audio signal 102 into the frames of the sequence of frames. In particular, DRC gains applicable to a particular frame of the sequence of frames may be inserted into that particular frame. Thus, each frame of the sequence of frames may include a DRC data component containing one or more explicit DRC gains to be applied to the respective frame. In particular, each frame may contain different explicit DRC gains for different rendering modes. For this purpose, DRC algorithms for different rendering modes may be applied within the encoder 150, and different DRC gains for different rendering modes may be determined within the encoder 150. The different DRC gains may then be explicitly inserted into the sequence of frames. As a result, the corresponding decoder 100 can directly apply explicit DRC gains without executing a DRC algorithm using a dynamic range compression curve.

[0096] Therefore, a sequence of frames may include, or indicate, multiple explicit DRC profiles for signaling dynamic range compression curves for multiple corresponding rendering modes. The multiple DRC profiles may be inserted into some (but not all) of the frames of the sequence of frames (e.g., frame I). Furthermore, a sequence of frames may include, or indicate, one or more DRC profiles for one or more corresponding rendering modes. Here, the one or more DRC profiles indicate that explicit DRC gains for one or more rendering modes are inserted into the frames of the sequence of frames. For example, the one or more DRC profiles for signaling explicit DRC gains may include a flag indicating whether the frames of the sequence of frames contain explicit DRC gains. The DRC gains may be inserted in each frame of the sequence of frames. In particular, each frame may contain the one or more DRC gains to be used to decode that frame.

[0097] The method may include inserting a DRC profile for an explicit DRC gain into a subset of frames from the sequence of frames. For example, the DRC profile on which the DRC gain for it is transmitted may indicate DRC configuration data for the explicit gain. In particular, the DRC profile on which the DRC gain for it is transmitted may be included in all of the subsets of DRC profiles. The DRC configuration data (e.g., a flag) may indicate that the sequence of frames contains an explicit DRC gain for a particular rendering mode. In this way, the decoder 100 is informed of the fact that for that particular rendering mode, the explicit DRC gain is derived directly from the frames of the sequence of frames.

[0098] Therefore, the method may further include determining an explicit DRC gain for the encoded audio signal 102 for a particular rendering mode. Furthermore, the method may include inserting the explicit DRC gain into the frames of the sequence of frames. The explicit DRC gain may be inserted into the frames from the sequence of frames to which the explicit DRC gain is applicable. Furthermore, the frames from the sequence of frames may include the one or more explicit DRC gains required to decode the frame within that particular rendering mode.

[0099] The method may further include inserting a DRC profile indicating DRC configuration data for the particular rendering mode into a subset of frames from the sequence of frames (for example, into I-frames). The DRC configuration data (including, for example, flags) may indicate that for that particular rendering mode, an explicit DRC gain is included within the frames of the sequence of frames. Thus, the decoder 100 can efficiently determine whether to use compression curves from multiple DRC profiles or explicit DRC gains to signal the dynamic range compression curve.

[0100] The DRC profile for signaling the dynamic range compression curve and the one or more DRC profiles pointing to an explicit DRC gain may be contained within a dedicated syntax element (for example, referred to as the DRC profile syntax element) of the I-frame of the sequence of frames.

[0101] The methods and systems described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or application-specific integrated circuits. Signals encountered in the methods and systems described may be stored on media such as random-access memory or optical storage media and transmitted over radio networks, satellite networks, wireless networks, or wired networks, such as the Internet. Typical devices utilizing the methods and systems described in this paper are portable electronic devices or other consumer equipment used to store and / or render audio signals.

[0102] Several aspects are described below. [Aspect 1] A method for generating an encoded audio signal, wherein the encoded audio signal has a sequence of frames, and the encoded audio signal exhibits multiple different dynamic range control (DRC) profiles for a plurality of different rendering modes, and the method is This includes inserting different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames so that two or more frames of the sequence of frames congruently contain the plurality of DRC profiles. method. [Aspect 2] The sequence of frames includes a subsequence consisting of I-frames; The different subsets of the DRC profile are inserted into different I-frames of the subsequence, which consists of I-frames. The method described in Embodiment 1. [Aspect 3] The method according to embodiment 1 or 2, wherein the subset of DRC profiles includes only a single DRC profile. [Aspect 4] The aforementioned DRC profiles include N DRC profiles, where N is an integer and N > 1; The N DRC profiles are inserted into N different frames from the sequence of frames. The method described in any one of the descriptions in 1 to 3. [Aspect 5] The method according to any one of embodiments 1 to 4, further comprising inserting all of the aforementioned DRC profiles into the first frame of the sequence of frames. [Aspect 6] The different subsets of DRC profiles from the plurality of DRC profiles are inserted into different frames of the sequence of frames such that each subsequence, consisting of M consecutive frames from the sequence of frames, congruently contains the plurality of DRC profiles; M is an integer, and M > 1. The method described in any one of the descriptions in paragraphs 1 to 5. [Aspect 7] The method according to any one of embodiments 1 to 6, further comprising inserting a flag into the frame of the sequence of frames, the flag indicating whether or not the frame contains a DRC profile. [Aspect 8] • One of the aforementioned DRC profiles is an explicit DRC profile that includes definitional data defining the dynamic range compression curve; The dynamic range compression curve provides a mapping between the input loudness and the gain that should be applied to the signal. The method described in any one of the descriptions in paragraphs 1 to 7. [Aspect 9] The method according to embodiment 8, wherein all of the aforementioned DRC profiles are explicit DRC profiles. [Aspect 10] The aforementioned definition data is: • Boost gain for boosting the aforementioned input loudness; • A boost gain range indicating the range of input loudness to which the boost gain can be applied; A null bandwidth indicating the range of the input loudness to which a gain of 0 dB can be applied; • Cut-off gain for attenuating the input loudness; • A cut-off gain range indicating the range of input loudness to which the cut-off gain is applicable; • A boost gain ratio that shows the transition between null gain and the boost gain; and / or • A cut-off gain ratio that shows the transition between the null gain and the cut-off gain. The method according to embodiment 8 or 9, comprising one or more of the above. [Aspect 11] The method according to any one of embodiments 8 to 10, further comprising inserting an implicit DRC profile instruction, wherein the definition data of the implicit DRC profile is known to the decoder of the encoded audio signal. [Aspect 12] The frames in the aforementioned sequence of frames include audio data and metadata; A subset of the DRC profile is inserted as metadata. The method described in any one of the descriptions in Actuals 1 to 11. [Aspect 13] A DRC profile includes definition data that defines the range of output reference levels to which the DRC profile is applicable; The aforementioned output reference level indicates the dynamic range of a certain rendering mode. The method described in any one of the descriptions in 1 to 12. [Aspect 14] The method according to embodiment 13, wherein the dynamic range of the rendering mode may decrease as the output reference level increases, and vice versa. [Aspect 15] The method according to embodiment 13 or 14, wherein the maximum boost gain and maximum cut gain of the dynamic range compression curve of the DRC profile may increase with increasing output reference level and vice versa. [Aspect 16] The aforementioned multiple DRC profiles are: • The first DRC profile for use in home theater rendering mode; • A second DRC profile for use in flat panel rendering mode; • A third DRC profile for use in portable device speaker rendering mode; and / or • Fourth DRC profile for use in headphone rendering mode A method according to any one of the embodiments 1 to 15, which includes one or more of the above. [Aspect 17] The method according to any one of embodiments 1 to 16, further comprising generating a bitstream containing the encoded audio signal, wherein the bitstream is an AC4 bitstream. [Aspect 18] - Determine the explicit DRC gain for the encoded audio signal for a specific rendering mode; The further includes inserting the explicit DRC gain into the frame of the sequence of frames, The method described in any one of the descriptions in Actuals 1 to 17. [Aspect 19] The method according to aspect 18, further comprising inserting a DRC profile having DRC configuration data for the particular rendering mode into a subset of frames of the sequence of frames, wherein the DRC configuration data indicates that for the particular rendering mode, there is an explicit DRC gain within the frames of the sequence of frames. [Aspect 20] • An explicit DRC gain is inserted into the frame from the sequence of frames to which the explicit DRC gain is applicable; and / or - A frame from the sequence of frames includes one or more explicit DRC gains required to decode that frame within that particular rendering mode. The method described in aspect 18 or 19. [Aspect 21] A bitstream comprising an encoded audio signal, wherein the encoded audio signal has a sequence of frames, the encoded audio signal represents a plurality of different dynamic range control (DRC) profiles for a plurality of different rendering modes, a plurality of subsets of DRC profiles from the plurality of DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently comprise the plurality of DRC profiles. [Aspect 22] A method for decoding an encoded audio signal, wherein the encoded audio signal has a sequence of frames, the encoded audio signal represents a plurality of different dynamic range control (DRC) profiles for a plurality of different rendering modes, a plurality of subsets of DRC profiles from the plurality of DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the plurality of DRC profiles, the method is The step of determining a first rendering mode from the aforementioned multiple different rendering modes; The steps include determining one or more DRC profiles from a subset of DRC profiles contained within the current frame of the sequence of frames; The step of determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode; If none of the one or more DRC profiles are applicable to the first rendering mode, the step is to select a default DRC profile as the current DRC profile, wherein the definition data of the default DRC profile is known in the decoder for decoding the encoded audio signal; This includes the step of decoding the current frame using the aforementioned DRC profile. method. [Aspect 23] The step (611) of determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode is - Determine the first output reference level for the first rendering mode; - Determine the range of output reference levels to which the DRC profiles from one or more of the above DRC profiles are applicable; - Includes determining whether the first output reference level falls within the range of output reference levels, The method described in aspect 22. [Aspect 24] The method according to embodiment 22 or 23, further comprising the step (604) of selecting a first DRC profile from the one or more DRC profiles as the current DRC profile if it is determined that the first DRC profile is applicable to the first rendering mode. [Aspect 25] The method according to any one of embodiments 22 to 24, further comprising the step of determining whether the current frame in the sequence of frames contains a subset of the DRC profile. [Aspect 26] A subset of the DRC profile is contained within the I-frame of the aforementioned sequence of frames; The step of determining whether the current frame contains a subset of the DRC profile includes determining whether the current frame is an I-frame (603), The method described in aspect 25. [Aspect 27] The step of determining whether the current frame contains a subset of DRC profiles includes verifying the DRC profile flags contained within the current frame (606), The method according to aspect 25 or 26. [Aspect 28] The current step involves determining whether the frame represents an implicit DRC profile from multiple implicit DRC profiles, wherein the definition data of the implicit DRC profile is known in the decoder that decodes the input audio signal; The step of selecting the implicit DRC profile as the current DRC profile if it is determined that the current frame represents an implicit DRC profile (608) The method described in any one of the descriptions in paragraphs 22 to 27. [Aspect 29] The method according to any one of embodiments 22 to 28, wherein the step of decoding the current frame includes leveling the sequence of frames to a first output reference level of the first rendering mode. [Aspect 30] The method according to any one of embodiments 22 to 29, wherein the step of decoding the current frame includes adapting the loudness level of the current frame using a dynamic range compression curve specified in the current DRC profile. [Aspect 31] An encoder for generating an encoded audio signal, wherein the encoded audio signal has a sequence of frames, and the encoded audio signal exhibits multiple different dynamic range control (DRC) profiles for multiple different rendering modes, and the encoder, The system is configured such that different subsets of DRC profiles from the plurality of DRC profiles are inserted into different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the plurality of DRC profiles. Encoder. [Aspect 32] A decoder for decoding an encoded audio signal, wherein the encoded audio signal has a sequence of frames, the encoded audio signal represents a plurality of different dynamic range control (DRC) profiles for a plurality of different rendering modes, a plurality of subsets of DRC profiles from the plurality of DRC profiles are contained within different frames of the sequence of frames, and two or more frames of the sequence of frames congruently contain the plurality of DRC profiles, and the decoder The step of determining a first rendering mode from the aforementioned multiple different rendering modes; The steps include determining one or more DRC profiles from a subset of DRC profiles contained within the current frame of the sequence of frames; The step of determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode; If none of the one or more DRC profiles are applicable to the first rendering mode, the step is to select a default DRC profile as the current DRC profile, wherein the definition data of the default DRC profile is known in the decoder; The system is configured to perform the step of decoding the current frame using the aforementioned DRC profile. decoder.

Claims

1. A method for decoding an encoded audio signal, wherein the encoded audio signal has a sequence of frames containing encoded audio data and metadata, the metadata includes a plurality of different sets of dynamic range control (DRC) gains and DRC configuration metadata in one or more frames of the sequence of frames, the DRC configuration metadata includes a plurality of DRC profiles associated with the encoded audio signal and, for each DRC profile, a range of output reference levels to which the DRC profile is applicable, each set of DRC gains corresponds to one of the plurality of DRC profiles, and the method is: - The step of setting the desired output reference level for the decoded audio signal; - The step of identifying one or more DRC profiles from among the DRC profiles, the range of applicable output reference levels that includes the desired output reference level for the decoded audio signal; - The step of selecting one of the identified DRC profiles; - The step of decoding the encoded audio signal; - The process includes the step of adjusting the dynamic range of the decoded audio signal by applying a DRC gain corresponding to the selected DRC profile to the decoded audio signal. The DRC gain corresponding to the selected DRC profile is equidistant in time. method.

2. A decoder for decoding an encoded audio signal, wherein the encoded audio signal has a sequence of frames containing encoded audio data and metadata, the metadata includes a plurality of different sets of dynamic range control (DRC) gains and DRC configuration metadata in one or more frames of the sequence of frames, the DRC configuration metadata includes a plurality of DRC profiles associated with the encoded audio signal and, for each DRC profile, a range of output reference levels to which that DRC profile is applicable, each set of DRC gains corresponds to one of the plurality of DRC profiles, and the decoder, - The step of setting the desired output reference level for the decoded audio signal; - The step of identifying one or more DRC profiles from among the DRC profiles, the range of applicable output reference levels that includes the desired output reference level for the decoded audio signal; - The step of selecting one of the identified DRC profiles; - The step of decoding the encoded audio signal; It has one or more processors that perform the step of adjusting the dynamic range of the decoded audio signal by applying a DRC gain corresponding to a selected DRC profile to the decoded audio signal. The DRC gain corresponding to the selected DRC profile is equidistant in time. decoder.

3. A non-temporary computer-readable storage medium having a sequence of instructions, wherein, when executed by an audio signal processing device, the sequence of instructions causes the audio signal processing device to execute the method described in claim 1.