Efficient DRC profile transmission
By inserting DRC profiles into audio frames, the method enables high-quality and intelligible audio playback across diverse devices and environments, addressing the challenge of varying dynamic range capabilities.
Patent Information
- Application Number
- JP2024172944
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-10-01
- Filing Date
- 2024-10-02
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2035-09-29
AI Technical Summary
Existing media playback systems face challenges in maintaining high-quality and intelligible audio playback across a wide range of devices with varying dynamic range capabilities and listening environments.
A method and system for transmitting Dynamic Range Control (DRC) profiles in a bandwidth-efficient manner by inserting different subsets of DRC profiles into audio frames, allowing decoders to select appropriate profiles for specific rendering modes.
Ensures high-quality and intelligible audio playback by adapting to different playback environments and devices, while reducing data overhead through efficient profile distribution.
Smart Images

Figure 0007821857000002 
Figure 0007821857000003 
Figure 0007821857000004
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 058,228, filed October 1, 2014, the contents of which are incorporated herein by reference in their entirety.
[0002] Technical Field This document relates to processing of audio signals. In particular, this document relates to a method and corresponding system for transmitting Dynamic Range Control (DRC) profiles in a bandwidth-efficient manner. [Background technology]
[0003] The growing popularity of media consumption devices has created new opportunities and challenges for creators and distributors of media content for playback on such devices, as well as for designers and manufacturers of such devices. Many consumer devices are capable of playing a wide range of media content types and formats, including those often associated with high-quality, wide-bandwidth, and wide-dynamic-range audio content for HDTV, Blu-ray, or DVD. Media processing devices can be used to play this type of audio content over their own internal acoustic transducers or over external transducers such as headphones or high-quality home theater systems. However, all these playback systems and environments impose significantly different demands on the dynamic range of the audio signal due to varying noise levels in the environment or the playback system's limited ability to reproduce the required sound pressure levels without distortion. Limiting the dynamic range depending on the environment is an approach to providing high quality and intelligibility across a wide range of different rendering devices with different rendering capabilities and listening environments, i.e., across a wide range of rendering modes. Summary of the Invention [Problem to be solved by the invention]
[0004] This paper addresses the technical challenges of media content creators and distributors with a bandwidth-efficient means to enable playback of audio signals with high quality and intelligibility on a wide range of different rendering devices with different rendering capabilities. [Means for solving the problem]
[0005] According to one aspect, a method is described for generating an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, the method including inserting different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames, such that two or more frames of the sequence of frames jointly include the plurality of DRC profiles.
[0006] According to a further aspect, a method for decoding an encoded audio signal is described. The encoded audio signal includes a sequence of frames. The encoded audio signal further includes a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes. Different subsets of DRC profiles from the plurality of DRC profiles are included in different frames of the sequence of frames, and two or more frames of the sequence of frames collectively include the plurality of DRC profiles. The method includes determining a first rendering mode from the plurality of different rendering modes and determining one or more DRC profiles from the subset of DRC profiles included in a current frame of the sequence of frames. The method further includes determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode. If none of the one or more DRC profiles is applicable to the first rendering mode, the method further includes selecting a default DRC profile as a current DRC profile. Here, definition data of the default DRC profile is known to a decoder for decoding the encoded audio signal. Additionally, the method includes decoding the current frame using the current DRC profile.
[0007] According to a further aspect, a bitstream including an encoded audio signal is described, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, the different subsets of DRC profiles from the plurality of DRC profiles being contained within different frames of the sequence of frames, and two or more frames of the sequence of frames jointly comprising the plurality of DRC profiles.
[0008] According to another aspect, an encoder for generating an encoded audio signal is described. The encoded audio signal has a sequence of frames. The encoded audio signal exhibits a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes. The encoder is configured to insert different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames, such that two or more frames of the sequence of frames jointly include the plurality of DRC profiles.
[0009] According to a further aspect, a decoder for decoding an encoded audio signal is described. The encoded audio signal includes a sequence of frames. The encoded audio signal exhibits a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes. The different subsets of DRC profiles from the plurality of DRC profiles are included in different frames of the sequence of frames, and two or more frames of the sequence of frames jointly include the plurality of DRC profiles. The decoder is configured to determine a first rendering mode from the plurality of different rendering modes, determine one or more DRC profiles from the subset of DRC profiles included in a current frame of the sequence of frames, determine whether at least one of the one or more DRC profiles is applicable to the first rendering mode, and select a default DRC profile as a current DRC profile if none of the one or more DRC profiles is applicable to the first rendering mode. Here, definition data of the default DRC profile is known to the decoder. The decoder is further configured to decode a current frame using the current DRC profile.
[0010] According to a further aspect, a software program is described that may be adapted for execution on a processor and, when executed on the processor, to perform the method steps outlined herein.
[0011] According to another aspect, a storage medium is described that may have a software program adapted for execution on a processor and, when executed on the processor, to perform the method steps outlined herein.
[0012] According to a further aspect, a computer program product is described, which may have executable instructions for performing the method steps outlined herein when executed on a computer.
[0013] It should be noted that the methods and systems, including the preferred embodiments, outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]
[0014] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 illustrates an exemplary audio decoder. [Figure 2] FIG. 1 illustrates an exemplary audio encoder. [Figure 3] FIG. 2 illustrates an exemplary dynamic range compression curve. [Figure 4] FIG. 2 illustrates an exemplary dynamic range compression curve. [Figure 5] FIG. 2 illustrates an exemplary sequence of frames. [Figure 6a]1 is a first half of a flowchart of an exemplary method for selecting a DRC profile. [Figure 6b] 10 is a flowchart illustrating a second half of an exemplary method for selecting a DRC profile. DETAILED DESCRIPTION OF THE INVENTION
[0015] As described above, this paper addresses the technical challenge of enabling audio content designers and / or distributors to control the quality and intelligibility of audio content for various types of rendering modes. An exemplary rendering mode is a home theater rendering mode, where audio content is played using transducers that typically allow a very wide dynamic range in a quiet environment. Another exemplary rendering mode is a flat panel mode, where audio content is played using transducers, such as those of a TV set, that typically allow a reduced dynamic range compared to a home theater. A further exemplary rendering mode is a portable speaker mode, where audio content is played using speakers of a portable electronic device (such as a smartphone). The dynamic range of this rendering mode is typically smaller than the rendering modes described above, and the environment is often noisy. Another exemplary rendering mode is a portable headphone mode, where audio content is played using headphones associated with the portable electronic device. The dynamic range is limited, but typically higher than the dynamic range provided by the speakers of the portable electronic device.
[0016] To allow high quality and high intelligibility for different rendering modes, different DRC (Dynamic Range Control) profiles for different rendering modes may be provided along with the audio content. The audio content may be transmitted in a sequence of frames. The sequence of frames may include I (i.e., independent) frames that can be decoded independently of preceding or following frames. Furthermore, the sequence of frames may include other types of frames (e.g., P and / or B frames) that typically exhibit dependency with respect to preceding and / or following frames. At least some frames in the sequence of frames may include multiple different DRC profiles for multiple different rendering modes. In particular, an I frame in the sequence of frames may include the multiple DRC profiles.
[0017] By inserting several different DRC profiles into a sequence of audio frames, the audio decoder can select the appropriate DRC profile for a particular rendering mode, ensuring that the rendered audio signal has high quality (in particular, no clipping or distortion introduced by the transducer) and high intelligibility.
[0018] Various aspects of dynamic range control are described below. Without customized dynamic range control, input audio information (e.g., PCM samples, time-frequency samples in a QMF matrix, etc.) is often reproduced in a playback device at a loudness level that is inappropriate for the specific playback environment of the playback device (i.e., including the physical and / or mechanical playback limitations of the device), which may differ from the target playback environment for which the encoded audio content was coded in the encoding device.
[0019] The techniques described in this document can be used to support dynamic range control of a wide variety of audio content customized for any of a wide variety of playback environments, while preserving the perceptual quality of the audio content and preserving the artist's intent to adapt the content to various playback environments.
[0020] Dynamic range control (DRC) refers to a time-varying, level-dependent audio processing operation that modifies a signal (e.g., compresses, cuts, stretches, boosts, etc.) to convert an input dynamic range of loudness levels in audio content into an output dynamic range that differs from the input dynamic range. For example, in one dynamic range control scenario, quiet sounds may be mapped (e.g., boosted) to higher loudness levels, and loud sounds may be mapped (e.g., cut) to lower loudness values. As a result, in the loudness domain, the output range of loudness levels is smaller than the input range of loudness levels in this example. However, in some embodiments, dynamic range control may be reversible, so that the original range can be restored. For example, a stretching operation can be performed to restore the original range, as long as the mapped loudness levels in the output dynamic range mapped from the original loudness levels are below the clipping level, each unique original loudness level is mapped to a unique output loudness level, etc.
[0021] The DRC techniques described herein can be used to provide a better listening experience in certain playback environments or situations. For example, soft sounds in a noisy environment may be masked by noise that makes the soft sounds inaudible. Conversely, loud sounds may be undesirable in some situations (e.g., in "late-night" listening modes), for example, because they may disturb neighbors. Many devices, typically with small form-factor speakers, cannot reproduce sound at high output levels or without perceptible distortion. In some cases, lower signal levels may be reproduced below the human hearing threshold. The DRC techniques may perform a mapping of input loudness levels to output loudness levels based on a DRC gain (e.g., a scaling factor such as scaling audio amplitude, boosting ratio, or cutting ratio) found using a dynamic range compression curve.
[0022] A dynamic range compression curve is a function (e.g., a look-up table, a curve, a multi-segment piecewise linear function) that maps individual input loudness levels (e.g., non-dialogue sounds) determined from individual audio data frames to corresponding output loudness levels, and thus to individual gain(s) for dynamic range control to convert the input loudness levels to the corresponding output loudness levels. Each individual gain indicates the amount of gain to be applied to a signal to map the corresponding individual input loudness level to the intended output loudness level. The output loudness level after applying the individual gains represents the target loudness level for the audio content in each audio data frame in a particular playback environment.
[0023] In addition to specifying a mapping between gain and loudness level, a dynamic range compression curve may include or be provided with specific release and attack times for applying a specific gain. Attack refers to an increase in signal energy (or loudness) between successive time samples, while release refers to a decrease in energy (or loudness) between successive time samples. The attack time (e.g., 10 ms, 20 ms, etc.) refers to the time constant used in smoothing the DRC gain when the corresponding signal is in attack mode. The release time (e.g., 80 ms, 100 ms, etc.) refers to the time constant used in smoothing the DRC gain when the corresponding signal is in release mode. In some embodiments, these time constants are additionally, optionally, or alternatively used to smooth the signal energy (or loudness) prior to determining the DRC gain.
[0024] Different dynamic range compression curves may correspond to different playback environments (i.e., different rendering modes). For example, a dynamic range compression curve for a playback environment of a flat-panel TV may be different from a dynamic range compression curve for a playback environment of a portable device. For example, a first dynamic range compression curve for a first playback environment of a portable device with speakers may be different from a second dynamic range compression curve for a second playback environment of the same portable device with a headset.
[0025] FIG. 1 shows a block diagram of exemplary components of an audio decoder 100. The audio decoder 100 includes a data extractor 104, a dynamic range controller 106, and an audio renderer 108. The data extractor 104 is configured to receive an encoded input signal 102. The encoded input signal 102 described herein may be a bitstream containing encoded (e.g., compressed) input audio data frames (particularly a sequence of audio frames) and possibly metadata. The bitstream may be an AC-4 bitstream. The data extractor 104 is configured to extract / decode the input audio data frames and metadata from the encoded input signal 102. Each input audio data frame includes multiple coded audio data blocks, each representing multiple audio samples. Each frame represents a (e.g., fixed) time interval containing a certain number of audio samples. The frame size may vary with the sample rate and coded data rate. An audio sample is a quantized audio data element (e.g., input PCM sample, input time-frequency sample in a QMF matrix, etc.) that represents the spectral content in one, two, or more (audio) frequency bands or frequency ranges. The quantized audio data elements in an input audio data frame may represent sound pressure waves in the digital (quantized) domain. The quantized audio data elements may cover a finite range of loudness levels up to a maximum possible value (e.g., clipping level, maximum loudness level, etc.).
[0026] The metadata can be used by the audio decoder 100 to process the input audio data frames. The metadata may include various operational parameters related to one or more operations to be performed by the decoder 100, one or more dynamic range compression curves (i.e., one or more DRC profiles), normalization parameters related to the dialogue loudness level represented in the input audio data frames, etc. Dialogue loudness level may refer to the dialogue loudness, program loudness, average dialogue loudness, etc. (e.g., psychoacoustic, perceptual, etc.) level of an entire program (e.g., a movie, television program, radio broadcast, etc.), a portion of a program, dialogue in a program, etc.
[0027] The operation and functionality of some or all of the decoder 100 or its modules (e.g., data extractor 104, dynamic range controller 106, etc.) may be adapted in response to metadata extracted from the encoded input signal 102. For example, the metadata—including, but not limited to, dynamic range compression curves, dialogue loudness levels, etc.—may be used by the decoder 100 to generate output audio data elements in the digital domain (e.g., output PCM samples, output time-frequency samples in a QMF matrix, etc.). The output data elements can then be used to drive audio channels or speakers to achieve a specified loudness or reference playback level during playback in a particular playback environment.
[0028] The dynamic range controller 106 may be configured to receive some or all of the audio data elements and metadata in the input audio data frames, and to perform audio processing operations (e.g., dynamic range control operations, gain smoothing operations, gain limiting operations, etc.) on the audio data elements in the input audio data frames based at least in part on the metadata extracted from the encoded audio signal 102.
[0029] In particular, the dynamic range controller 106 may include a selector 110, a loudness calculator 112, and a DRC gain unit 114. The selector 110 may be configured to determine a speaker configuration associated with a particular playback environment in the decoder 100 (e.g., home theater mode, flat panel mode, portable device mode with speakers, portable device mode with headphones, 5.1 speaker configuration mode, 7.1 speaker configuration mode, etc.). Furthermore, the selector 110 may be configured to select a particular dynamic range compression curve (i.e., a DRC profile) from dynamic range compression curves extracted from metadata of the encoded input signal 102 (i.e., from the plurality of DRC profiles).
[0030] The loudness calculator 112 may be configured to calculate one or more types of loudness levels represented by audio data elements in the input audio data frame. Examples of the types of loudness levels include, but are not limited to, any of: individual loudness levels across individual frequency bands in individual channels across individual time intervals; broadband loudness levels across a wide frequency range in individual channels; loudness levels determined from or smoothed across an audio data block or frame; loudness levels determined from or smoothed across two or more audio data blocks or frames; loudness levels smoothed across one or more time intervals; etc. Zero, one, or more of these loudness levels may be modified by the decoder 100 for dynamic range control.
[0031] To determine the loudness level, the loudness calculator 112 may determine one or more time-dependent physical wave attributes, such as spatial and / or local pressure levels at specific audio frequencies, represented by audio data elements in the input audio data frame. The loudness calculator 112 may use the one or more time-varying physical wave attributes to derive one or more types of loudness levels based on one or more psychoacoustic functions that model human loudness perception. The psychoacoustic functions may be nonlinear functions—constructed based on a model of the human auditory system—that convert / map specific spatial pressure levels at specific audio frequencies to specific loudness for the specific audio frequencies.
[0032] A loudness level across multiple (audio) frequencies or multiple frequency bands (e.g., broadband, wideband, etc.) may be derived through integration of specific loudness levels across multiple (audio) frequencies or multiple frequency bands. A time-averaged, smoothed, etc. loudness level across one or more time intervals (e.g., longer than represented by the audio data elements in an audio data block or frame) may be obtained using one or more smoothing filters implemented as part of the audio processing operations in decoder 100. Another exemplary method for determining a (broadband) loudness level is specified in ITU-R BS.1770, which applies time-domain filtering to the time-domain input audio signal, then calculates the RMS (root-mean-square) level for each channel of the input audio signal, and then integrates and gates the resulting loudness levels across the channels.
[0033] Specific loudness levels for different frequency bands may be calculated for each audio data block of a certain (e.g., 256) samples. A pre-filter may be used to apply frequency weighting (e.g., similar to IEC B weighting) to the specific loudness levels before integrating them into a broadband (or wideband) loudness level. A summation of broad loudness levels across two or more channels (e.g., left front, right front, center, left surround, right surround, etc.) may be performed to provide an overall loudness level for the two or more channels.
[0034] An overall loudness level may refer to the broadband loudness level of a single channel (e.g., center) of a speaker configuration. An overall loudness level may refer to the broadband loudness level of multiple channels. The multiple channels may be all channels of a speaker configuration (i.e., for a rendering mode). Additionally, optionally, or alternatively, the multiple channels may include a subset of channels of a speaker configuration (e.g., a subset of channels including left front, right front, and low frequency effects (LFE); a subset of channels including left surround and right surround; a subset of channels including center, etc.).
[0035] The loudness level (e.g., broadband, wideband, global, specific, etc.) may be used as input to find a corresponding DRC gain (e.g., static, pre-smoothing, pre-limiting, etc.) from a selected dynamic range compression curve. The loudness level used as input to find the DRC gain may first be adjusted or normalized with respect to a dialogue loudness level from metadata extracted from the encoded audio signal 102 and / or with respect to an output reference level of a rendering mode. The adjustment and normalization relative to the dialogue loudness level / output reference level may be performed on a portion of the audio signal in the encoded audio signal 102 in a non-loudness domain (e.g., SPL domain) before a specific spatial pressure level represented in that portion of the audio content in the encoded audio signal 102 is converted to a specific loudness level for that portion of the audio content in the encoded audio signal 102.
[0036] The DRC gain unit 114 may be configured with a DRC algorithm to generate a gain (e.g., for dynamic range control, gain limiting, gain smoothing, etc.) and apply the gain to one or more loudness levels at one or more types of loudness levels represented by audio data elements in the input audio data frame to achieve a target loudness level for the particular playback environment. The application of a gain (e.g., a DRC gain) as described herein may occur in the loudness domain. For example, the gain may be generated based on a loudness calculation (which may be expressed in terms of sones or simply SPL compensated for, e.g., dialogue loudness levels without conversion), smoothed, and applied directly to the input signal. Techniques as described herein may apply a gain to a signal in the loudness domain, then calculate a corresponding gain to be applied to the signal by converting the signal from the loudness domain back to the (linear) SPL domain, and evaluating the signal in the loudness domain before and after the gain has been applied to the signal. The ratio (or difference when expressed in log-dB notation) then determines the corresponding gain for that signal.
[0037] The DRC algorithm may operate with multiple DRC parameters, including a dialogue loudness level that has already been calculated and embedded in the encoded audio signal 102 by an upstream encoder 150 (e.g., as described in the context of FIG. 2 ) and can be obtained by the decoder 100 from metadata in the encoded audio signal 102. The dialogue loudness level from the upstream encoder 150 indicates an average dialogue loudness level (e.g., per program, relative to the energy of a full-scale 1 kHz sine wave, relative to the energy of a reference square wave, etc.). The dialogue loudness level extracted from the encoded audio signal 102 may be used to reduce loudness level differences between programs. The reference dialogue loudness level may be set to the same value for different programs in the same specific playback environment in the decoder 100. Based on the dialogue loudness level from the metadata, the DRC gain unit 114 can apply a dialogue loudness-related gain to each audio data block in the program so that the output dialogue loudness level (or output reference level) averaged across multiple audio data blocks of the program is raised or lowered to a (e.g., pre-configured, system default, user-configurable, profile-dependent, etc.) reference dialogue loudness level for that program. The dialogue loudness level may be used to calibrate the DRC algorithm. In particular, the null band of the DRC algorithm may be adjusted to the dialogue loudness level. Alternatively, the desired output reference level may be used to calibrate the DRC algorithm when it is applied to a signal to which gain has been applied to change the dialogue loudness level to equal the desired output reference level. The dialogue loudness level may correspond to a so-called dialnorm parameter. This is the case when speech gating is applied to determine the dialnorm parameter.In some embodiments, the dialogue loudness level corresponds to a dialnorm parameter determined by gating based on a loudness level threshold rather than by using speech gating.
[0038] DRC gains may be used to address differences in loudness levels within a program by boosting or cutting signal portions that are soft and / or loud according to a selected dynamic range compression curve. One or more of these DRC gains may be calculated / determined by a DRC algorithm based on the selected dynamic range compression curve and the loudness level (e.g., broadband, wideband, global, specific, etc.) determined from one or more corresponding audio data blocks, audio data frames, etc.
[0039] The loudness level used to determine the DRC gain (e.g., static, pre-smoothing, pre-gain-limiting, etc.) by searching a selected dynamic range compression curve may be calculated over a short interval (e.g., approximately 5.3 milliseconds). The integration time of the human auditory system can be much longer (e.g., approximately 200 milliseconds). The DRC gain obtained from the selected dynamic range compression curve may be smoothed with a time constant to account for the long integration time of the human auditory system. To implement a fast rate of change (increase or decrease) in loudness level, a short time constant may be used to cause the loudness level to change over a short time interval corresponding to the short time constant. Conversely, to implement a slow rate of change (increase or decrease) in loudness level, a long time constant may be used to change the loudness level over a long time interval corresponding to the long time constant.
[0040] The human auditory system may respond to increasing and decreasing loudness levels with different integration times. To smooth the static DRC gain retrieved from the selected dynamic range compression curve, different time constants may be used depending on whether the loudness level is increasing or decreasing. For example, to match the characteristics of the human auditory system, attacks (increases in loudness level) may be smoothed with a relatively short time constant (e.g., attack time), while releases (decrease in loudness level) may be smoothed with a relatively long time constant (e.g., release time).
[0041] A DRC gain for a portion of audio content (e.g., one or more audio data blocks, audio data frames, etc.) may be calculated using a loudness level determined from the portion of audio content. The loudness level to be used for lookup in the selected dynamic range compression curve may first be adjusted with respect to (e.g., relative to) a dialogue loudness level (e.g., of a program of which the audio content is a part) in metadata extracted from the encoded audio signal 102.
[0042] Reference Dialogue Loudness Level / Output Reference Level (e.g. -31dB in "Line" mode) FS , -20dB in "RF" mode FS ) may be specified or established for a particular playback environment in the decoder 100. Additionally, alternatively or optionally, in some embodiments, a user may be given control over setting or changing a reference dialogue loudness level in the decoder 100.
[0043] The DRC gain unit 114 may be configured to determine a dialogue loudness related gain for the audio content to cause a change from an input dialogue loudness level to a reference dialogue loudness level as the output dialogue loudness level.
[0044] The audio renderer 108 may be configured to generate channel-specific (e.g., multi-channel) audio data 116 for the particular speaker configuration after applying a gain determined based on DRC, gain limiting, gain smoothing, etc. to the input audio data extracted from the encoded audio signal 102. The channel-specific audio data 116 may be used to drive speakers, headphones, etc. represented in the speaker configuration.
[0045] Additionally and / or optionally, the decoder 100 may be configured to perform one or more other operations relating to processing, rendering, downmixing, resampling, etc. relating to the input audio data.
[0046] The techniques described herein can be used with a variety of speaker configurations corresponding to a variety of different surround sound configurations (e.g., 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10-60 speaker configurations, 60+ speaker configurations, object signals or combinations of object signals, etc.) and a variety of different rendering environment configurations (e.g., movie theaters, parks, opera houses, concert halls, bars, homes, auditoriums, etc.).
[0047] 2 shows an exemplary encoder 150. Encoder 150 may include an audio content interface 152, a dialogue loudness analyzer 154, a DRC reference repository 156, and an audio signal encoder 158. Encoder 150 may be part of a broadcast system, an internet-based content server, an over-the-air network operator system, a film production system, etc.
[0048] Audio content interface 152 may be configured to receive audio content 160 and audio content control input 162 and generate encoded audio signal 102 based at least in part on some or all of audio content 160 and audio content control input 162. For example, audio content interface 152 may be used to receive audio content 160 and audio content control input 162 from a content creator, content provider, etc.
[0049] Audio Content 160 may be part of or comprise all of an overall media data set, including audio only, audiovisual, etc. Audio Content 160 may include one or more of program portions, a program, several programs, one or more commercials, etc.
[0050] The dialogue loudness analyzer 154 may be configured to determine / establish one or more dialogue loudness levels for one or more portions of the audio content 152 (e.g., one or more programs, one or more commercials, etc.). The audio content may be represented by one or more collections of audio tracks. Dialogue audio content of the audio content may be on a separate audio track, and / or at least a portion of the dialogue audio content of the audio content may be on an audio track containing non-dialogue audio content.
[0051] Audio content control inputs 162 may include some or all of user control inputs, control inputs provided by systems / devices external to encoder 150, control inputs from a content creator, control inputs from a content provider, etc. For example, a user, such as a mixing engineer, may provide / specify one or more dynamic range compression curve identifiers, which may be used to retrieve one or more dynamic range compression curves that best fit the audio content 160 from a data repository, such as the DRC reference repository (156).
[0052] The DRC reference storage 156 may be configured to store DRC reference parameter sets, etc. These DRC reference parameter sets may include definition data for one or more dynamic range compression curves, etc. The encoder 150 may encode (e.g., simultaneously) two or more dynamic range compression curves into the encoded audio signal 102. Zero, one, or more of the dynamic range compression curves may be standard-based, proprietary, customized, decoder-modifiable, etc. As an example, the dynamic range compression curves of FIGS. 3 and 4 may be encoded (e.g., simultaneously) into the encoded audio signal 102.
[0053] The audio signal encoder 158 may be configured to receive audio content from the audio content interface 152, a dialogue loudness level from the dialogue loudness analyzer 154, retrieve one or more DRC reference parameter sets (i.e., DRC profiles) from the DRC reference store 156, format the audio content into audio data blocks / frames, format the dialogue loudness levels, DRC reference parameter sets, etc. into metadata (e.g., metadata containers, metadata fields, metadata structures, etc.), and encode the audio data blocks / frames and metadata into the encoded audio signal 102. Audio content to be encoded into an encoded audio signal as described herein may be received in one or more of a variety of source audio formats in one or more of a variety of ways, such as wirelessly, via a wired connection, through a file, via internet download, etc.
[0054] The encoded audio signal 102 described herein can be part of an overall media data bitstream (e.g., for an audio broadcast, an audio program, an audiovisual program, an audiovisual broadcast, etc.). The media data bitstream can be accessed from a server, a computer, a media storage device, a media database, a media file, etc. The media data bitstream can be broadcast, transmitted, or received over one or more wireless or wired network links. The media data bitstream can be communicated through one or more intermediaries, such as a network connection, a USB connection, a wide area network, a local area network, a radio connection, an optical connection, a bus, a crossbar connection, a serial connection, etc.
[0055] Any of the components depicted (e.g., in Figures 1 and 2) may be implemented as one or more processes and / or one or more integrated circuit circuits (e.g., ASICs, FPGAs, etc.) in hardware, software, or a combination of hardware and software.
[0056] 3 and 4 show exemplary dynamic range compression curves that can be used by the DRC gain unit 104 in the decoder 100 to derive a DRC gain from an input loudness level. As shown, the dynamic range compression curve may be centered around a reference loudness level in the program (e.g., an output reference level) to provide an overall gain appropriate for a particular playback environment. Exemplary definition data for the dynamic range compression curve (e.g., in the metadata of the encoded audio signal 102) (e.g., including, but not limited to, boost ratios, cut ratios, attack times, release times, etc.) are shown in the table below. Different profiles (e.g., film standard, film light, music standard, music light, speech, etc.) represent different playback environments (e.g., in the decoder 100).
[0057] [Table 1] dB SPL or dB FS Loudness level expressed in dB and SPL One or more compression curves may be received that are described using gains expressed in dB relative to the SPL A different loudness representation (e.g., sones) that has a nonlinear relationship with loudness level may be implemented, and the compression curve used in the DRC gain calculation may then be transformed to be described using the different loudness representation (e.g., sones).
[0058] 5 shows an exemplary encoded audio signal 102 including a sequence of frames (numbered n+1 through n+30, where n is an integer). In the illustrated example, every fifth frame is an I-frame. In the illustrated example, I-frame (n+1) has multiple DRC profiles (identified as home theater AVR (audio / video receiver), flat panel, portable HP (headphones), and portable SP (speaker)). Each DRC profile has a dynamic range compression curve as shown in FIGS. 3 and 4.
[0059] The multiple DRC profiles may be repeatedly inserted within an I-frame of a sequence of frames. This allows the decoder 100 to determine an appropriate DRC profile for the encoded audio signal 102 and for the current rendering mode at startup of the decoded audio signal 102, when tuning in to an audio program on air, and / or after a splice point. On the other hand, repeatedly transmitting a complete set of DRC profiles leads to a relatively high bitstream overhead. In view of this, it is proposed to transmit a varying subset of DRC profiles within an I-frame of the encoded audio signal 102.
[0060] 5 shows an example for inserting a DRC profile within a sequence of frames. In the illustrated example, only a single DRC profile from the complete set of DRC profiles is inserted into an I-frame. The DRC profile inserted into an I-frame changes for each I-frame, so that after N I-frames (N=4 in the illustrated example), the decoder 100 has received the complete set of N DRC profiles. This reduces the data rate for transmitting the complete set of DRC profiles while ensuring that the decoder 100 receives the complete set of DRC profiles within a reasonable amount of time.
[0061] 6a and 6b show a flowchart of an exemplary method 600 for determining a DRC profile for decoding a frame of the encoded audio signal 102. The method 600 may be performed by the decoder 100 (particularly the selector 110). At the start of reception of the encoded audio signal 102, the DRC profile used by the decoder 100 may be initialized. The DRC profile used to decode the current frame of the encoded audio signal 102 may be referred to as the current DRC profile. Thus, at startup, the current DRC profile may be initialized. In particular, a default DRC profile (available in the decoder 100) may be set to be the current DRC profile used to render the current frame (method step 601). Thus, the variable "profile" may be set to the default DRC profile (profile=defaultDRCprofile). Furthermore, the decoder 100 may keep track of a previously used profile. The previously used profile may be set to undefined (prev_profile=undefined).
[0062] The method 600 may further include step 602 of fetching a new frame to be decoded (i.e., a current frame) from the encoded audio signal 102. In step 603, it is verified whether the new frame is an I-frame, which may include a DRC profile. If the new frame is not an I-frame, the method 600 proceeds to step 604, where the current DRC profile is used to process the new frame. Further, in method step 605, the previously used profile is set to the current DRC profile (prev_profile=profile).
[0063] If the new frame is an I-frame, the I-frame may be checked to see if it contains DRC data at method step 606. By way of example, the metadata of the I-frame may include a flag indicating whether the I-frame contains DRC data. If DRC data is not present, method 300 may proceed to steps 604, 605. Otherwise, the method may proceed to method step 607.
[0064] In method step 607, it may be verified whether the new frame is the first frame of the encoded audio signal 102 to be decoded. As can be seen from the flowcharts of FIGS. 6a and 6b, this may be verified by checking the prev_profile variable. If the prev_profile variable is undefined, the new frame is the first frame to be decoded. If the new frame is the first frame to be decoded, the decoder 100 may use a predefined DRC profile other than the default DRC profile. For this purpose, the metadata of the new frame may include an identifier (ID) for such a predefined DRC profile. Such a predefined DRC profile may be stored in a database at the decoder 100. The use of a predefined DRC profile may provide a bitrate-efficient means for signaling the DRC profile to be used to the decoder 100, since only the ID of the predefined profile needs to be transmitted (method step 608). Predefined DRC profiles signaled using an ID may be referred to as implicit DRC profiles.
[0065] In some cases, it may be beneficial to only use a single predefined DRC profile other than the default DRC profile, in which case the decoder 100 may be configured to set the profile variable to that predefined (i.e., implicit) DRC profile without receiving any ID in the metadata of the new frame.
[0066] Method 600 may further include verifying whether the metadata of the new frame includes one or more explicit DRC profiles (step 609). An explicit DRC profile includes an ID for identifying the explicit DRC profile. Furthermore, an explicit DRC profile typically includes definition data for a dynamic range compression curve, such as those shown in FIGS. 3 and 4. The dynamic range compression curve may be defined as a piecewise linear function. Furthermore, an explicit DRC profile may indicate a range of output reference levels (ORLs) to which the explicit DRC profile is applicable. As an example, a default DRC profile and / or the predefined (implicit) DRC profile may be applicable for output reference levels ranging from −31 dB FS to 0 dB FS.
[0067] The ORL of a rendering device may indicate the dynamic range capability of the rendering device. Typically, the dynamic range capability decreases as the ORL increases. For a high ORL, a compression curve with a high degree of compression should be used to render the audio signal in an intelligible way without clipping. On the other hand, for a low ORL, the compression may be reduced to render the audio signal with a high dynamic range. Due to the high dynamic range capability of the rendering device, the intelligibility of the audio signal is still guaranteed.
[0068] If the metadata of the new frame includes at least one explicit DRC profile, the profile data of the first DRC profile is read (step 610). Furthermore, it is verified whether the ORL range of the first DRC profile is applicable to the currently used rendering device (step 611). If not, the method 600 proceeds to search for another explicit DRC profile in the metadata of the new frame. On the other hand, if an explicit DRC profile is applicable to the rendering device, this explicit DRC profile may be set as the current DRC profile to be used to process the new frame (step 614).
[0069] Method 600 may further include verifying whether a headphone rendering mode is used and whether an explicit DRC profile is applicable to the headphone rendering mode (step 612). Method 600 may further include verifying whether the explicit DRC profile is an updated profile compared to a previously used profile (step 613). For this purpose, the ID of the explicit DRC profile may be compared with the ID of the currently used profile. This ensures that decoder 100 always uses the latest current DRC profile.
[0070] Using method 600, it may be ensured that decoder 100 always identifies a DRC profile for rendering a frame of encoded audio signal 102, even if the data has not yet received a DRC profile for the current rendering mode (i.e., for the current rendering device). Furthermore, it is ensured that the DRC profile for the current rendering mode is applied as soon as decoder 100 receives the corresponding DRC profile.
[0071] Thus, a method 600 for decoding an encoded audio signal 102 is described. The encoded audio signal 102 comprises a sequence of frames. Furthermore, the encoded audio signal 102 exhibits a plurality of different dynamic range control (DRC) profiles for corresponding a plurality of different rendering modes. Examples of different rendering modes (or different playback environments) are a first DRC profile for use in a home theater rendering mode; a second DRC profile for use in a flat panel rendering mode; a third DRC profile for use in a portable device speaker rendering mode; and / or a fourth DRC profile for use in a headphone rendering mode. The DRC profile defines a specific DRC behavior. The DRC behavior may be described by a compression curve (and time constant) and / or by a DRC gain. The DRC gains may be temporally equidistant gains that can be applied to the encoded audio signal 102 to implement a DRC. The compression curve may be accompanied by time constants that together constitute a DRC algorithm. DRC typically reduces the volume of loud sounds and amplifies quiet sounds, thereby compressing the dynamic range of the audio signal for an improved experience in non-ideal playback environments.
[0072] A sequence of frames typically includes multiple consecutive frames that make up an audio signal. An audio program (e.g., a broadcast television or radio program) may include multiple audio signals that are joined at splice points. For example, a main audio program may be repeatedly interrupted by commercial breaks. A sequence of frames may correspond to a complete audio program. Alternatively, a sequence of frames may correspond to one of the multiple audio signals that make up the complete audio program.
[0073] Different subsets of DRC profiles from the plurality of DRC profiles may be included in different frames of a sequence of frames, such that two or more frames of the sequence of frames jointly include the plurality of DRC profiles. As noted above, distribution of DRC profiles across multiple frames of a sequence of frames results in reduced bitstream overhead for signaling the plurality of DRC profiles.
[0074] The method 600 may include determining a first rendering mode from the plurality of different rendering modes. In particular, which rendering mode is used to render the encoded audio signal 102 may be determined. Furthermore, the method 600 may include determining 609, 610 one or more DRC profiles from the plurality of DRC profiles included in a current frame of the sequence of frames. In other words, one or more DRC profiles from a subset of the DRC profiles included in the current frame may be determined. Furthermore, it may be determined 611 whether at least one of the one or more DRC profiles is applicable to the first rendering mode. Determining 611 whether at least one of the one or more DRC profiles is applicable to the first rendering mode may include determining a first output reference level for the first rendering mode, determining a range of output reference levels within which a DRC profile from the one or more DRC profiles is applicable, and determining whether the first output reference level is within the range of output reference levels.
[0075] The method 600 may further include selecting 604 a default DRC profile as a current DRC profile if none of the one or more DRC profiles is applicable to the first rendering mode. Definition data of the default DRC profile is typically known to a decoder for decoding the encoded audio signal 102. The method 600 may further include decoding (and / or rendering) the current frame using the current DRC profile. Thus, it may be ensured that the decoder 100 utilizes a DRC profile (and dynamic range compression curve) even if the decoder 100 has not yet received a DRC profile specific to the encoded audio signal 102.
[0076] Alternatively or additionally, the method 600 may include selecting 604 a first DRC profile from the one or more DRC profiles as a current DRC profile if it is determined that the first DRC profile is applicable to the first rendering mode, such that the decoder 100 is configured to use the first DRC profile that is optimal for the encoded audio signal 102 and for the first rendering mode as soon as the decoder 100 receives the first DRC profile.
[0077] The method 600 may further include determining 603, 606 whether a current frame of the sequence of frames includes one or more DRC profiles from the plurality of DRC profiles, i.e., whether the current frame includes a subset of DRC profiles. As outlined in the context of FIG. 5 , a subset of DRC profiles is typically included in an I-frame of the sequence of frames. Thus, determining 603, 606 whether the current frame includes one or more DRC profiles from the plurality of DRC profiles or whether the current frame includes a subset of DRC profiles may include determining 603 whether the current frame is an I-frame. As noted above, an I-frame may be a frame that is decodable independently of any other frame from the sequence of frames. This may be due to the fact that data included in such an I-frame is transmitted in a manner that does not depend on data from previous or subsequent frames. In particular, data in an I-frame is not differentially encoded with respect to data included in previous or subsequent frames.
[0078] Furthermore, determining 603, 606 whether the current frame includes one or more DRC profiles from the plurality of DRC profiles, or whether the current frame includes a subset of DRC profiles, may include verifying a DRC profile flag included in the current frame 606. The DRC profile flag in the bitstream of the encoded audio signal provides a bandwidth- and computationally-efficient means for identifying frames that carry a DRC profile.
[0079] The method 600 may further include determining whether the current frame indicates an implicit DRC profile from multiple implicit DRC profiles. The implicit DRC profile may include predefined legacy compression curves and time constants that can be used to transcode to E-AC-3. As described above, definition data for the implicit DRC profile may be known in the decoder 100 for decoding the input audio signal 102. In contrast to a default DRC profile, an implicit DRC profile may be specific to various types of audio signals (e.g., as shown in Table 1). The current frame of the sequence of frames may indicate a specific implicit DRC profile (e.g., using an identifier ID). This may provide a bandwidth-efficient means for signaling an appropriate DRC profile for the encoded audio signal 102. The implicit DRC profile may be selected 608 as the current DRC profile if it is determined that the current frame indicates the implicit DRC profile.
[0080] Decoding the current frame may include leveling the sequence of frames to the first output reference level of the first rendering mode. Further, decoding the current frame may include adapting a loudness level of the current frame using a dynamic range compression curve specified in the current DRC profile. The loudness level adaptation may be performed as outlined in the context of FIG. 1.
[0081] Depending on the number of frames from the sequence of frames, the current DRC profile may correspond to a default DRC profile (which is typically independent of the input audio signal 102), an implicit DRC profile (which may be adapted to the input audio signal 102 in a limited way) or the first explicit DRC profile (which may be designed for the input audio signal 102 and / or the first rendering mode).
[0082] Typically, only a subset of frames includes a DRC profile. Once a current DRC profile is selected, the current DRC profile may be maintained for decoding frames of the sequence of frames that do not include any DRC profile. Furthermore, even if a frame with a DRC profile is received, the current DRC profile may be maintained unless a DRC profile is received that is newer and / or more relevant to the encoded audio signal 102 than the current DRC profile (where the first selected explicit DRC profile is more relevant than the selected implicit DRC profile, which is more relevant than the default DRC profile). This ensures continuity and optimality of the DRC profiles used.
[0083] As a complement to the method 600 for decoding an encoded audio signal 102, a method for generating or encoding an encoded audio signal 102 is described. The encoded audio signal 102 comprises a sequence of frames. Furthermore, the encoded audio signal 102 exhibits a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes. The method includes inserting different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames, such that two or more frames of the sequence of frames jointly include the plurality of DRC profiles. In other words, subsets of DRC profiles that are less than the total number of DRC profiles may be provided with various frames of the sequence of frames. This may reduce the overhead of the encoded audio signal 102 while still providing a complete set of DRC profiles to a corresponding decoder 100. In other words, an advantage of this approach is that the encoder 150 has increased flexibility in how to transmit DRC data. This flexibility can be used to reduce the bit rate.
[0084] A sequence of frames may include a sub-sequence of I-frames (e.g., every Xth frame of the sequence of frames may be an I-frame). Different subsets of DRC profiles may be inserted into different (e.g., consecutive) I-frames of the sub-sequence of I-frames. To further reduce bandwidth, I-frames may be skipped, i.e., some I-frames may not contain DRC profile data.
[0085] A subset (e.g., each subset) of DRC profiles may include only a single DRC profile. In particular, the plurality of DRC profiles may include N DRC profiles, where N is an integer and N>1. The N DRC profiles may be inserted into N different frames from the sequence of frames. In this way, the bit rate required for transmission of the DRC profiles may be minimized.
[0086] The method may further include inserting all of the plurality of DRC profiles into a first frame of the sequence of frames (e.g., into the first frame of the sequence of frames of the audio signal). As a result, rendering of the encoded audio signal 102 may begin directly with the correct explicit DRC profile. As mentioned above, an audio program may be divided into multiple partial audio programs. For example, a main audio program may be interrupted by a commercial break. It may be beneficial to insert all of the plurality of DRC profiles into the first frame of each partial audio program. In other words, it may be beneficial to insert all of the plurality of DRC profiles immediately after the one or more junction points of an audio program including multiple partial audio programs.
[0087] Various subsets of DRC profiles from the plurality of DRC profiles may be inserted into different frames of the sequence of frames, such that each of M directly consecutive subsequences from the sequence of frames jointly comprises the plurality of DRC profiles, where M is an integer and M>1. In other words, the plurality of DRC profiles may be repeatedly transmitted within a block of M frames. As a result, the decoder 100 needs to wait at most M frames before obtaining an optimal explicit DRC profile for the encoded audio signal 102.
[0088] The method may further include inserting a flag into a frame of the sequence of frames, where the flag indicates whether the frame includes a DRC profile. By providing such a flag, a corresponding decoder 100 can efficiently identify frames that include DRC profile data.
[0089] A DRC profile of the plurality of DRC profiles may be an explicit DRC profile that includes (ie, carries) definition data for defining a dynamic range compression curve.
[0090] As outlined herein, the dynamic range compression curve provides a mapping between input loudness and output loudness and / or a gain to be applied to an audio signal. In particular, the definition data may include one or more of: a boost gain for boosting the input loudness; a boost gain range indicating a range for the input loudness to which the boost gain is applicable; a null band range indicating a range of input loudness to which a gain of 0 dB is applicable; a cut gain for attenuating the input loudness; a cut gain range indicating a range of input loudness to which the cut gain is applicable; a boost gain ratio indicating a transition between a null gain and the boost gain; and / or a cut gain ratio indicating a transition between the null gain and the cut gain.
[0091] The method may further include inserting an indication (e.g., an identifier, ID) of an implicit DRC profile, where definition data of the implicit DRC profile is typically known to the decoder 100 of the encoded audio signal 102. The implicit DRC profile indication may provide a bandwidth-efficient means for signaling a DRC profile that is (in a limited way) adapted to the encoded audio signal 102.
[0092] As outlined above, the frames of the sequence of frames typically contain audio data and metadata, and a subset of the DRC profile is typically inserted as metadata.
[0093] A DRC profile may include definition data that defines a range of output reference levels for which the DRC profile is applicable. The output reference level typically indicates the dynamic range of a given rendering mode. In particular, the dynamic range of a rendering mode may decrease with increasing output reference level, or vice versa. Furthermore, the maximum boost gain and maximum cut gain of the dynamic range compression curve of the DRC profile may increase with increasing output reference level, or vice versa. Thus, the output reference level provides an efficient means for selecting an appropriate DRC profile (with an appropriate dynamic range compression curve) for a particular rendering mode.
[0094] The method may further include generating a bitstream including the encoded audio signal 102. The bitstream may be an AC4 bitstream, i.e., the bitstream may conform to the AC4 bitstream format.
[0095] The method may further include inserting explicit DRC gains for the encoded audio signal 102 into frames of the sequence of frames. In particular, a DRC gain applicable to a particular frame of the sequence of frames may be inserted into the particular frame. Thus, each frame of the sequence of frames may include a DRC data component including one or more explicit DRC gains to be applied to the respective frame. In particular, each frame may include different explicit DRC gains for different rendering modes. For this purpose, DRC algorithms for different rendering modes may be applied within the encoder 150, and different DRC gains for the different rendering modes may be determined in the encoder 150. The different DRC gains may then be explicitly inserted into the sequence of frames. As a result, the corresponding decoder 100 can directly apply explicit DRC gains without executing a DRC algorithm using a dynamic range compression curve.
[0096] Thus, a sequence of frames may include or indicate multiple explicit DRC profiles to signal dynamic range compression curves for multiple corresponding rendering modes. The multiple DRC profiles may be inserted into some (but not all) of the frames (e.g., I-frames) of the sequence of frames. Furthermore, a sequence of frames may include or indicate one or more DRC profiles for one or more corresponding rendering modes, where the one or more DRC profiles indicate that explicit DRC gains for one or more rendering modes are inserted into the frames of the sequence of frames. By way of example, the one or more DRC profiles for signaling explicit DRC gains may include a flag indicating whether the frames of the sequence of frames include explicit DRC gains. The DRC gains may be inserted into each frame of the sequence of frames. In particular, each frame may include the one or more DRC gains to be used to decode that frame.
[0097] The method may include inserting a DRC profile for an explicit DRC gain into a subset of frames from the sequence of frames. By way of example, the DRC profile for which the DRC gain is transmitted may indicate DRC configuration data for the explicit gain. In particular, the DRC profile for which the DRC gain is transmitted may be included in all of the subset of DRC profiles. The DRC configuration data (e.g., a flag) may indicate that the sequence of frames includes explicit DRC gains for a particular rendering mode. In this way, the decoder 100 is informed about the fact that, for that particular rendering mode, explicit DRC gains are derived directly from frames of the sequence of frames.
[0098] Thus, the method may further include determining explicit DRC gains for the encoded audio signal 102 for a particular rendering mode. Furthermore, the method may include inserting the explicit DRC gains into frames of the sequence of frames. Explicit DRC gains may be inserted into frames from the sequence of frames to which the explicit DRC gains are applicable. Furthermore, a frame from the sequence of frames may include the one or more explicit DRC gains required to decode the frame in the particular rendering mode.
[0099] The method may further include inserting a DRC profile into a subset of frames from the sequence of frames (e.g., into an I-frame) that indicates DRC configuration data for the particular rendering mode. The DRC configuration data (e.g., including a flag) may indicate the fact that an explicit DRC gain is included in a frame of the sequence of frames for the particular rendering mode. Thus, decoder 100 may efficiently determine whether to use compression curves from multiple DRC profiles or explicit DRC gains to signal a dynamic range compression curve.
[0100] The DRC profile for signaling a dynamic range compression curve and the one or more DRC profiles pointing to an explicit DRC gain may be included within a dedicated syntax element (e.g., referred to as a DRC profile syntax element) of an I frame of the sequence of frames.
[0101] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or application-specific integrated circuits. Signals encountered in the described methods and systems may be stored on media such as random access memory or optical storage media, or transmitted over radio, satellite, wireless, or wired networks, such as the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer equipment used to store and / or render audio signals.
[0102] Several aspects will be described. [Aspect 1] 1. A method for generating an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, the method comprising: inserting different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames such that two or more frames of the sequence of frames jointly include the plurality of DRC profiles; method. [Aspect 2] the sequence of frames includes a sub-sequence of I-frames; the different subsets of DRC profiles are inserted into different I-frames of the subsequence of I-frames, The method of embodiment 1. Aspect 3 3. The method of embodiment 1 or 2, wherein the subset of DRC profiles comprises only a single DRC profile. Aspect 4 the plurality of DRC profiles includes N DRC profiles, where N is an integer and N>1; the N DRC profiles are inserted into N different frames from the sequence of frames; 4. The method of any one of embodiments 1 to 3. Aspect 5 5. The method of any one of aspects 1-4, further comprising inserting all of the plurality of DRC profiles into a first frame of the sequence of frames. Aspect 6 the different subsets of DRC profiles from the plurality of DRC profiles are inserted into different frames of the sequence of frames such that each subsequence of M consecutive frames from the sequence of frames jointly comprises the plurality of DRC profiles; M is an integer and M>1; 6. The method of any one of embodiments 1 to 5. Aspect 7 7. The method of any one of aspects 1-6, further comprising inserting a flag into a frame of the sequence of frames, the flag indicating whether the frame includes a DRC profile. Aspect 8 a DRC profile of the plurality of DRC profiles is an explicit DRC profile including definition data defining a dynamic range compression curve; The dynamic range compression curve gives the mapping between the input loudness and the gain that should be applied to the signal. 8. The method of any one of embodiments 1 to 7. Aspect 9 9. The method of embodiment 8, wherein all of the plurality of DRC profiles are explicit DRC profiles. Aspect 10 The definition data is: a boost gain for boosting the input loudness; a boost gain range indicating a range for the input loudness over which the boost gain is applicable; Null band range indicating the range of input loudness for which a gain of 0 dB is applicable; Cut gain for attenuating the input loudness; a cut gain range indicating the range of the input loudness for which the cut gain is applicable; a boost gain ratio indicating the transition between the null gain and the boost gain; and / or a cut gain ratio indicating the transition between the null gain and the cut gain; 10. The method of embodiment 8 or 9, comprising one or more of: Aspect 11 11. The method of any one of aspects 8 to 10, further comprising inserting an indication of an implicit DRC profile, wherein definition data of the implicit DRC profile is known to a decoder of the encoded audio signal. Aspect 12 a frame of said sequence of frames comprising audio data and metadata; A subset of the DRC profile is inserted as metadata, 12. The method of any one of embodiments 1 to 11. Aspect 13 The DRC profile includes definition data defining a range of power reference levels to which the DRC profile is applicable; The output reference level indicates the dynamic range of a certain rendering mode. 13. The method of any one of embodiments 1 to 12. Aspect 14 14. The method of claim 13, wherein the dynamic range of the rendering mode may decrease with increasing output reference level, and vice versa. Aspect 15 15. The method of embodiment 13 or 14, wherein the maximum boost gain and maximum cut gain of the dynamic range compression curve of the DRC profile may increase with increasing output reference level, or vice versa. Aspect 16 The plurality of DRC profiles: First DRC profile for use in Home Theater rendering mode; A second DRC profile for use in flat panel rendering mode; A third DRC profile for use in portable device speaker rendering mode; and / or A fourth DRC profile for use in headphone rendering mode 16. The method of any one of embodiments 1 to 15, comprising one or more of: Aspect 17 17. The method of any one of aspects 1-16, further comprising generating a bitstream including the encoded audio signal, wherein the bitstream is an AC4 bitstream. Aspect 18 determining an explicit DRC gain for the encoded audio signal for a particular rendering mode; further comprising inserting the explicit DRC gain into a frame of the sequence of frames. 18. The method of any one of embodiments 1 to 17. Aspect 19 The method of aspect 18, further comprising inserting a DRC profile having DRC configuration data for the particular rendering mode into a subset of frames of the sequence of frames, the DRC configuration data indicating the fact that an explicit DRC gain is included in a frame of the sequence of frames for the particular rendering mode. Aspect 20 An explicit DRC gain is inserted into a frame from the sequence of frames for which the explicit DRC gain is applicable; and / or a frame from said sequence of frames includes said one or more explicit DRC gains required to decode that frame within that particular rendering mode, 20. The method of embodiment 18 or 19. Aspect 21 1. A bitstream including an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, different subsets of DRC profiles from the plurality of DRC profiles being included in different frames of the sequence of frames, and two or more frames of the sequence of frames jointly including the plurality of DRC profiles. Aspect 22 1. A method of decoding an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, different subsets of DRC profiles from the plurality of DRC profiles being included in different frames of the sequence of frames, and two or more frames of the sequence of frames jointly comprising the plurality of DRC profiles, the method comprising: determining a first rendering mode from the plurality of different rendering modes; determining one or more DRC profiles from a subset of DRC profiles included in a current frame of said sequence of frames; determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode; if none of the one or more DRC profiles is applicable to the first rendering mode, selecting a default DRC profile as a current DRC profile, wherein definition data of the default DRC profile is known in a decoder for decoding the encoded audio signal; decoding a current frame using the current DRC profile; method. Aspect 23 determining (611) whether at least one of the one or more DRC profiles is applicable to the first rendering mode; determining a first output reference level for the first rendering mode; determining a range of power reference levels to which a DRC profile from the one or more DRC profiles is applicable; determining whether the first output reference level is within the range of output reference levels; 23. The method of embodiment 22. Aspect 24 The method of embodiment 22 or 23, further comprising a step (604) of selecting a first DRC profile from the one or more DRC profiles as a current DRC profile if it is determined that the first DRC profile is applicable to the first rendering mode. Aspect 25 25. The method of any one of embodiments 22-24, further comprising determining whether a current frame of the sequence of frames includes a subset of a DRC profile. Aspect 26 the subset of the DRC profile is contained within an I-frame of said sequence of frames; determining whether the current frame includes a subset of the DRC profile includes determining whether the current frame is an I-frame (603); 26. The method of embodiment 25. Aspect 27 determining whether the current frame includes a subset of the DRC profile includes verifying (606) a DRC profile flag included in the current frame; 27. The method of embodiment 25 or 26. Aspect 28 determining whether a current frame indicates an implicit DRC profile from a plurality of implicit DRC profiles, where definition data of the implicit DRC profile is known in a decoder that decodes the input audio signal; If it is determined that the current frame indicates an implicit DRC profile, selecting (608) the implicit DRC profile as the current DRC profile; 28. The method of any one of embodiments 22 to 27. Aspect 29 29. The method of any one of aspects 22 to 28, wherein the decoding of the current frame includes leveling the sequence of frames to a first output reference level of the first rendering mode. Aspect 30 30. The method of any one of aspects 22 to 29, wherein the decoding of the current frame includes adapting the loudness level of the current frame using a dynamic range compression curve specified in the current DRC profile. Aspect 31 1. An encoder for generating an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, the encoder comprising: configured to insert different subsets of DRC profiles from the plurality of DRC profiles into different frames of the sequence of frames, such that two or more frames of the sequence of frames jointly comprise the plurality of DRC profiles; Encoder. Aspect 32 1. A decoder for decoding an encoded audio signal, the encoded audio signal having a sequence of frames, the encoded audio signal exhibiting a plurality of different dynamic range control (DRC) profiles for a corresponding plurality of different rendering modes, different subsets of DRC profiles from the plurality of DRC profiles being included in different frames of the sequence of frames, and two or more frames of the sequence of frames jointly comprising the plurality of DRC profiles, the decoder comprising: determining a first rendering mode from the plurality of different rendering modes; determining one or more DRC profiles from a subset of DRC profiles included in a current frame of said sequence of frames; determining whether at least one of the one or more DRC profiles is applicable to the first rendering mode; if none of the one or more DRC profiles is applicable to the first rendering mode, selecting a default DRC profile as a current DRC profile, wherein definition data of the default DRC profile is known in the decoder; and decoding the current frame using the current DRC profile. decoder.
Claims
1. 1. A method of decoding an encoded audio signal, the encoded audio signal having a sequence of frames including encoded audio data and metadata, the metadata including a plurality of different sets of Dynamic Range Control (referred to as DRC) gains and DRC configuration metadata for one or more frames of the sequence of frames, the DRC configuration metadata indicating a plurality of DRC profiles associated with the encoded audio signal and, for each DRC profile, a range of Output Reference Levels for which the DRC profile is applicable, each set of DRC gains corresponding to one of the plurality of DRC profiles, the method comprising: - setting a desired output reference level for the decoded audio signal; - identifying one or more of the DRC profiles whose applicable output reference level range includes the desired output reference level for a decoded audio signal; selecting one of the identified DRC profiles; - decoding the encoded audio signal; adjusting the dynamic range of the decoded audio signal by applying a DRC gain corresponding to the selected DRC profile to the decoded audio signal; method.
2. The encoded audio signal further comprises a measure of broadband loudness of the audio signal, the method further comprising: determining a loudness-related gain in response to the measure of broadband loudness of the audio signal and the desired output reference level for the decoded audio signal; applying the loudness-related gain to the adjusted decoded audio signal to obtain a loudness-adjusted decoded audio signal having the desired output reference level. The method of claim 1.
3. 1. A decoder for decoding an encoded audio signal, the encoded audio signal having a sequence of frames including encoded audio data and metadata, the metadata including a plurality of different sets of dynamic range control (referred to as DRC) gains and DRC configuration metadata for one or more frames of the sequence of frames, the DRC configuration metadata indicating a plurality of DRC profiles associated with the encoded audio signal and, for each DRC profile, a range of output reference levels for which the DRC profile is applicable, each set of DRC gains corresponding to one of the plurality of DRC profiles, the decoder comprising: - setting a desired output reference level for the decoded audio signal; - identifying one or more of the DRC profiles whose applicable output reference level range includes the desired output reference level for a decoded audio signal; selecting one of the identified DRC profiles; - decoding the encoded audio signal; adjusting the dynamic range of the decoded audio signal by applying a DRC gain corresponding to the selected DRC profile to the decoded audio signal; decoder.
4. The encoded audio signal further comprises a measure of broadband loudness of the audio signal, and the one or more processors further: determining a loudness-related gain in response to the measure of broadband loudness of the audio signal and the desired output reference level for the decoded audio signal; applying the loudness-related gain to the adjusted decoded audio signal to obtain a loudness-adjusted decoded audio signal having the desired output reference level; 4. The decoder of claim 3, wherein the decoder executes:
5. 10. A non-transitory computer-readable storage medium having a sequence of instructions that, when executed by an audio signal processing device, causes the audio signal processing device to perform the method of claim 1.
Citation Information
Patent Citations
System and method for optimizing loudness and dynamic range across different playback devices
WO2014113471A1