Deferred loudness adjustment for dynamic range control
By deferring loudness normalization to the decoder side, the patent addresses inaccuracies in metadata-based DRC for live audio, enhancing the listening experience through accurate dynamic range control.
Patent Information
- Application Number
- JP2024059054
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-10
- Filing Date
- 2024-04-01
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-11-10
AI Technical Summary
Metadata-based dynamic range control (DRC) for live audio streaming and recording is challenging when the program loudness is not yet known, leading to inaccurate compressor characteristics and undesirable loudness shifts.
Postpone DRC loudness adjustment from the encoder side to the decoder side, applying loudness normalization using integrated loudness updates and smoothing filters to ensure accurate dynamic range control during playback.
Improves listening experience by reducing undesirable loudness shifts and pumping effects, especially in live streaming and recording applications.
Smart Images

Figure 0007815308000002 
Figure 0007815308000003 
Figure 0007815308000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio decoder device, and in particular to deferred loudness adjustment for dynamic range control. [Background technology]
[0002] A sound program, such as music, a podcast, a short video clip recorded live, or a feature-length film, has loud and soft parts that define its dynamics and dynamic range. In many situations, such as listening through a headset in a noisy environment or listening through loudspeakers at home late at night, it is desirable to reduce the dynamics and dynamic range of the reproduced sound to improve the listener's experience. For this purpose, a dynamic range compressor is used. This compressor is a digital signal processor that applies a time-varying gain to the input, which is a digital audio signal (of a sound program), to amplify the soft parts of the audio signal and attenuate the loud parts. To avoid audible pumping artifacts that may result from compressing the dynamic range of an audio signal, a loudness normalization process can be performed that "matches" the input audio signal to a compression profile or profile while compressing the audio signal according to the compression profile. This can be done by offsetting the instantaneous loudness of an input audio signal by the program loudness of that signal, which is a calculated value that aims to represent the overall loudness of a sound program (also called integrated loudness). Summary of the Invention
[0003] Audio coding standards define dynamic range compression methods that generate dynamic range control (DRC) gains at the encoder side, where a sound program is created or prepared for distribution or storage / archive. In this specification, DRC gain refers to a DRC gain sequence that is time-aligned with an associated sound program, such that one or more gain values in the sequence are applied to corresponding digital audio frames in the sound program. The DRC gain sequence is then formatted into one or more bitstreams, for example, as metadata associated with the sound program. The decoder side obtains the bitstreams and, when desired at the decoder side (typically during playback of the decoded audio signal), applies the DRC gains in the stream to compress the dynamic range of the decoded audio signal. The advantage of such a metadata-based approach is improved quality, since a longer look-ahead time interval can be obtained for offline encoding of the DRC gains than is possible with real-time compression. Another advantage is that compression characteristics can be controlled at the encoder side, for example, by the expertise of the sound program creator or distributor.
[0004] Metadata-based DRC in online applications (e.g., live audio streaming and recording of live audio to a file) presents a challenge when the program loudness of a sound program being streamed for playback or written for storage is not yet known (because the sound program has not yet finished), because the compressor characteristics may not be properly adjusted (or loudness normalized) if the actual program loudness of the sound program (which can only be determined once the sound program has finished) deviates significantly from what is expected or predicted.
[0005] Some aspects of the present disclosure are novel digital signal processing methods that postpone dynamic range control (DRC) loudness adjustment (loudness normalization) from the encoder side to the decoder side. Another aspect is a technique for modifying compressor characteristics at the decoder side when metadata-based DRC gain sequencing is used for loudness normalization. These aspects are particularly useful for applications such as live streaming and live recording to a file.
[0006] The above summary does not provide an exhaustive list of all aspects of the present disclosure. The present disclosure is believed to include all systems and methods possible from all suitable combinations of the various aspects summarized above, as well as those disclosed in the following detailed description and particularly pointed out in the claims. Such combinations may have particular advantages not specifically recited in the above summary.
[0007] Several aspects of the present disclosure are illustrated herein by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference characters indicate like elements. It should be noted that references to "an" or "one" aspect of the present disclosure are not necessarily to the same aspect, but rather they mean at least one. Also, for the sake of brevity and reducing the total number of figures, a given figure may be used to illustrate features of multiple aspects of the present disclosure, and not all elements in a figure may be required for a given aspect. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an exemplary DRC characteristic curve. [Figure 2] FIG. 1 is a block diagram of an audio codec system that applies DRC on the decoder side and does not perform loudness normalization on the encoder side. [Figure 3]FIG. 1 is a block diagram of an audio codec system suitable for live streaming, where DRC is applied on the decoder side and loudness normalization is not performed on the encoder side. [Figure 4] FIG. 1 is a block diagram of an audio codec system suitable for live recording to storage or archive, with DRC applied at the decoder side and no loudness normalization at the encoder side. [Figure 5] FIG. 1 illustrates a portion of an MPEG-D DRC compliant audio codec system that applies DRC at the decoder side. [Figure 6] FIG. 1 illustrates a portion of an MPEG-D DRC compliant audio codec system with loudness normalization on the encoder side and DRC applied on the decoder side. [Figure 7] 1 shows part of an MPEG-D DRC compliant audio codec system that applies DRC with loudness normalization at the decoder side. [Figure 8] FIG. 1 is a flow diagram of a new encoder-side process that can generate backwards-compatible and non-backwards-compatible MPEG-D DRC bitstream extensions. [Figure 9] FIG. 10 is a flow diagram of a new decoder-side process that can generate a DRC gain sequence using either backward compatible or non-backward compatible MPEG-D DRC bitstream extensions. [Figure 10A] FIG. 1 is a block diagram of an MPEG-D DRC compliant audio codec system in which a backward compatible encoder produces a backward compatible bitstream that is processed by both new and legacy decoders. [Figure 10B] FIG. 1 is a block diagram of an MPEG-D DRC compliant audio codec system in which a backward compatible encoder produces a backward compatible bitstream that is processed by both new and legacy decoders. DETAILED DESCRIPTION OF THE INVENTION
[0009] Some aspects of the present disclosure will now be described with reference to the accompanying drawings. Whenever the shape, relative position, and other aspects of the described parts are not explicitly defined, it is meant that the scope of the present disclosure is not limited to only the parts shown, which is for illustrative purposes only. Also, while numerous details are described, it will be understood that some aspects of the present disclosure can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail in order to avoid obscuring an understanding of this specification.
[0010] To properly apply dynamic range control to an audio signal, the compressor characteristics (DRC characteristics, DRC profile) should be "matched" to the loudness level range of the audio signal. For example, referring to FIG. 1, the matching is performed along the input level axis so that the zero crossing of the DRC characteristic curve is approximately at the center of the loudness level range of the audio signal. The level at the zero crossing point is also referred to as the DRC input loudness target, and in the exemplary set of characteristic curves shown in FIG. 1, that level is approximately -31 dB. The center of the loudness level range may be, for example, the average level of the sound program or the average dialogue level within the sound program. In this specification, the process of achieving such a match is referred to as loudness normalization to a given loudness target associated with the DRC of the audio signal. For example, the loudness of an audio signal (sound program) may be a single value known as integrated loudness. Integrated loudness is a measure of the loudness of an audio signal, similar to root mean square (RMS) but with higher fidelity from the perspective of human hearing. Integrated loudness may be equivalent to program loudness in that it measures how loud a sound program is over its entire duration. To achieve loudness normalization, if integrated loudness is given in units of decibels (dB), it can be subtracted from the DRC input loudness target to derive a normalization gain in dB. This normalization gain is added to the output of a loudness model that calculates the instantaneous loudness of the audio signal (sound program). The instantaneous loudness may be a sequence of loudness values (representing human perceived loudness) calculated based on each digital audio frame that makes up the input digital audio signal. Another way to achieve loudness normalization is to shift the DRC characteristic curve shown in Figure 1 to the right or left (by the amount of normalization gain).In the example of Figure 1, the curve has been shifted to the left by -31 dB (the loudness target in this example) and is therefore well matched to (and can be directly applied to) sound programs with an integrated loudness of -31 dBA (A-weighted) or LKFS (loudness K-weighted level full scale). In other words, the normalization gain in that case is zero dBA.
[0011] If the integrated loudness of the audio program is not yet known during dynamic range control signal processing, as is the case with live audio, a prediction must be made in order to apply loudness normalization. However, the prediction may be inaccurate, resulting in DRC gains that contain undesirable biases or that produce a pumping effect, which is an undesirable loudness shift between the uncompressed and compressed portions of the audio signal.
[0012] To reduce the possibility of undesired loudness shifts, one aspect of the present disclosure applies loudness normalization to the DRC at the decoder side rather than the encoder side of the audio codec system or method. An example of an audio codec system and related method is shown in the hardware block diagram of FIG. 2. Various hardware blocks of the audio codec system and method may be implemented by a programmed processor. In such a method, the integrated loudness (necessary for loudness normalization performed in conjunction with the DRC for playback or archiving / storage of the decoded audio signal) can be obtained in at least two examples described below in conjunction with FIGS. 3 and 4.
[0013] Referring first to FIG. 2, an audio codec system has an encoder side, which may be implemented, for example, in one or more servers, by one or more processors, collectively referred to herein as "programmed processors," that execute or are configured by instructions stored in memory. The upper audio signal processing path includes a side chain that includes a loudness model that calculates or estimates the instantaneous loudness of a digital input audio signal (sequence), also referred to herein as a sound program. This estimation is based on a perceptual loudness scale (such as the Sohn scale) and is therefore approximately logarithmic. To smooth the instantaneous loudness sequence over time, a smoothing filter can be applied, as shown. As a result, regions of the input audio sequence where compression gain changes are undesired are smoothed, while macro-dynamic loudness transitions are unaffected.
[0014] The lower audio signal processing path includes a delay block that delays the input audio sequence to compensate for the side chain delay. The smoothed loudness sequence and the delayed input audio sequence are then fed to the encoder.
[0015] The encoder can perform bitrate reduction operations on one or both inputs to generate one or more bitstreams containing bitrate-reduced versions of one or both inputs. The one or more bitstreams can then be transmitted to the decoder side (e.g., via the Internet) or written to a file for storage or archival until accessed by decoder-side processing. The smoothed loudness (referring to a sequence of smoothed loudness values or a single smoothed loudness value) may be carried as metadata in the same bitstream as the delayed input audio sequence, e.g., associated with a "corresponding" Advanced Audio Coding (AAC) audio frame. This is also referred to as being present in the audio layer. Alternatively, the smoothed loudness sequence and other loudness values, such as the aggregate loudness update value and DRC payload (as discussed further below), may be transmitted at a higher layer, such as the file format level, rather than in the audio layer. In either case, one or more bitstreams are generated in which the encoded audio is provided along with associated metadata, such as the smoothed loudness sequence, or instructions for the decoder side to apply the DRC gain sequence provided by the encoder, as otherwise described below.
[0016] The decoder side may also be implemented as a programmed processor, such as one or more processors executed by or configured by instructions stored in a memory as part of an audio playback device. Note that the decoder-side processing may be implemented in the same audio playback device as the encoder-side processing. Alternatively, the decoder-side processing may be implemented in an audio playback device separate from the programmed processor performing the encoder-side processing. Examples of audio playback devices include smartphones, tablet computers, digital media players, headsets, or vehicle infotainment systems. On the decoder side, the decoder removes the encoder's bitrate reduction operation to recover the smoothed loudness sequence and the delayed input audio sequence. Once decoded, the smoothed loudness values are then mapped to "corresponding" DRC or compression gain values. This mapping is a memoryless input-output function that implements, for example, one of the curves shown in FIG. 1 (or any other desired curve). This mapping forms a compressor characteristic or compressor profile (DRC characteristic), the output of which is a time-varying gain (a sequence of DRC gain values) that is a function of the time-varying input loudness level. This mapping may also include a conversion from the logarithmic loudness domain to the linear domain of DRC gains. If compression is desired, the DRC gain value (sequence) is then applied to the decoded audio signal, as indicated by the multiplication symbol in the diagram. Although not shown, the compressed audio is then passed to a playback processing block, which may ultimately generate a transducer (speaker) driver input signal that converts the compressed audio into sound.
[0017] 2, it can be seen that the smoothed loudness sequence is adjusted or normalized at the decoder side before being input to the DRC mapping block. For example, the unchanged integrated loudness (a single value) can be subtracted (in the dB domain) from the DRC input loudness target to derive a normalization gain in dB. This normalization gain is added to each of the smoothed loudness values in the smoothed loudness sequence to generate a normalized loudness sequence used in the DRC processing. There are at least two application examples for such DRC processing: live or real-time streaming, and live recording to a file for storage or archiving.
[0018] In one such application, referring to FIG. 3, the input audio on the encoder side is a live or real-time digital audio recording being streamed to the decoder side, for example, via the Internet. The input audio includes an audio capture of a live or real-time event, simultaneously with encoding and bitstream transmission. Therefore, a single integrated loudness value representing the entire sound program cannot be calculated until the live event ends. In the meantime, an integrated loudness measurement block within the encoder side collects samples of the live audio, delayed for time alignment and sent to the encoder, over a time interval longer than a single audio frame, e.g., several seconds, and calculates a loudness measurement for that time interval. This block then "integrates" or collects several such measurements, moving backwards through the beginning of the sound program, and, for example, averages them to calculate an integrated loudness update value. The integrated loudness update value may be a measurement of the integrated loudness for only the portion of the sound program that has been played or streamed up to the current update. For example, periodically, this measurement is repeated to generate a "running average" integrated loudness, and the latest integrated loudness update value (a single value) is sent to the decoder side. It should be noted that the term "running average," as used herein, does not require an actual averaging, but only some measurement of the loudness of a sound program from the beginning of the program to the current update, based on a collection of loudness measurements, including a statistical evaluation of the collected loudness measurements. The update (running average) may be calculated and then provided as part of a bitstream that also includes the encoded sound program (encoded audio signal) as multiple instances of the aggregate loudness update field, with adjacent instances in the bitstream separated by 1 to 10 seconds over the duration of the sound program.
[0019] It should also be noted that the term "integrated loudness update" may also be referred to as running average loudness or "partial integrated loudness." At the end of a sound program, the last or final loudness update may represent the loudness of the entire sound program (e.g., also referred to as integrated loudness or program loudness, as described in Recommendation ITU-R BS.1770-4 (10 / 2015) algorithm for measuring audio program loudness and true peak audio levels).
[0020] On the decoder side, the decoder takes the bitstream, extracts the integrated loudness update value from it, and then the decoder-side processing applies it to the DRC processing to perform loudness normalization. This may be done, for example, by adding a single loudness normalization gain value (e.g., the difference between the DRC input loudness target and the integrated loudness update value) to the decoded or restored instantaneous loudness sequence before inputting it into the DRC feature mapping block. Alternatively, loudness normalization may be performed by shifting the DRC feature along the input axis by an amount equal to the loudness normalization gain value. The loudness normalization gain may be periodically updated during the transmission of the bitstream (sound program) using the latest partial integrated loudness value (integrated loudness update value) calculated at the encoder side for the elapsed part of the live event.
[0021] In another application example, referring to FIG. 4, the input audio on the encoder side is a live or real-time digital audio recording of an event, which is written to a file for archival or storage purposes at the end of the recording (when the event ends). A single integrated loudness value representing the program loudness of the entire live audio event can be calculated by the integrated loudness model block at the end of the recording and provided to the encoder as soon as the event ends. The encoder writes the integrated loudness value to a file, along with an encoded version of the live audio and an encoded version of the instantaneous (and smoothed) loudness sequence calculated by the loudness model (based on the same live audio). On the decoder side, the decoder obtains a file (bitstream), decodes the input audio and instantaneous loudness sequence from the file, and extracts the integrated loudness value from the file. The decoder-side processing then applies loudness normalization to the decoded instantaneous loudness sequence using the integrated loudness value before inputting it to the DRC (compression) mapping block, and the output of this block is then applied to the decoded input audio during playback (if compression is desired).
[0022] In one embodiment, the smoothing filter is a nonlinear filter such as that described in U.S. Patent No. 10,109,288. A useful property of this filter is that its output can be level-shifted by the same amount as its input. That is, if we define f(x) as the nonlinear function, x(n) as the input signal, and y(n) as the output, we can write: y(n)=f(x(n))
[0023] If a shift in the input signal is given by ΔL, and the output is shifted by ΔL, then f(x) satisfies the shift property, which can be expressed mathematically as follows: y(n)+ΔL=f(x(n)+ΔL)
[0024] This is beneficial as it avoids any encoder-side side-chain processing that has a dependency on the absolute loudness value.
[0025] Another aspect of the present disclosure is a method for applying DRC in accordance with the MPEG-D DRC standard ISO / IEC, "Information technology—MPEG Audio Technologies—Part 4: Dynamic Range Control," ISO / IEC 23003-4:2020 ("MPEG-D DRC"), extended to accommodate loudness normalization at the encoder side. Figure 5 shows a simplified block diagram of a portion of the MPEG-D DRC process, where DRC gains are generated and applied based on decoding the DRC gains from metadata in the bitstream obtained from the encoder side. MPEG-D DRC provides predefined DRC characteristics and a flexible way to encode parameterized characteristics.
[0026] In Figure 5, the encoder side applies the smoothed instantaneous loudness sequence (calculated for the input audio sequence) to a selected DRC characteristic (also referred to as a "mapping block" as used above in connection with Figure 2). The output of the DRC characteristic mapping block generates a DRC gain sequence, which is provided to a DRC encoder. The DRC encoder performs bit-rate reduction and encodes the input sequence into one or more bitstreams, which are then transmitted or otherwise made available to the decoder side. At the decoder side, a DRC decoder removes the bit-rate reduction encoding to recover the DRC gain sequence (the decoded DRC gain sequence). The decoded DRC gain sequence is then applied to the decoded audio signal (if compression is desired).
[0027] MPEG-D DRC also supports a type of decoder-side processing that changes the DRC characteristics applied to compress a sound program from those used at the encoder side (to calculate the DRC gain sequence inserted into the bitstream as metadata) as shown in FIG. 5 to another that may be selected by the decoder-side processing based on the current playback or listening conditions. This is achieved by first applying the encoder-side DRC gain sequence to inverse characteristic A as shown in FIG. 6. Inverse characteristic A is the inverse of DRC characteristic A applied at the encoder side to generate the encoder-side DRC gain sequence. An index (identifier or pointer) to DRC characteristic A (used at the encoder side to generate the DRC gain sequence) may be provided in the bitstream so that the decoder side can identify the inverse characteristic A. Applying the DRC gain sequence as input to inverse characteristic A results in a restored smoothed instantaneous loudness sequence. Ignoring quantization effects, the restored loudness sequence (at the output of the inverse characteristic A block) is essentially the smoothed loudness sequence used by the encoder-side processing. As a result, the restored loudness sequence can be applied to a second DRC characteristic B to generate a second DRC gain sequence that may be better suited (than DRC characteristic A) for compressing the decoded audio signal. The second DRC gain sequence is then applied to the decoded audio (e.g., if compression is desired during playback).
[0028] According to one aspect of the present disclosure, the loudness normalization in the encoder-side side chain shown in FIG. 6 is replaced using the approach shown in FIG. 2. That is, an offset (normalization gain) based on the integrated loudness is applied at the decoder side instead of the encoder side. FIG. 7 shows a block diagram of such a system. Herein, this system is also referred to as an extended MPEG-D DRC-compliant system (hereinafter also referred to as having a "new" encoder and a "new" decoder). Such a system has a block called Integrated Loudness Measurement at the encoder side, the output of which provides an integrated loudness update value as discussed above with respect to FIG. 3. This integrated loudness update value is provided to an audio encoder. Herein, this encoder is a DRC encoder that also encodes a DRC gain sequence (in addition to the input audio). The DRC gain sequence may be determined as discussed above with respect to FIG. 6. The encoded DRC gain sequence and the integrated loudness update value are provided to the decoder side via one or more bitstreams. The DRC gain sequence may be formatted as metadata associated with the coded input audio, which is also provided to the decoder side.
[0029] The integrated loudness measurement is a moving measurement of integrated loudness (also referred to herein as a moving average) that begins taking at the beginning of a sound program and continues over time to "integrate" the audio signal of the sound program in order to calculate the integrated loudness for only the elapsed portion of the sound program. As the audio signal (sound program) continues, the integrated loudness measurement generates updates, for example, periodically, for example every 10 seconds. These integrated loudness updates are written to the bitstream (for example, by a DRC encoder). In MPEG-D DRC, this can be achieved either by writing the updates to an extension field or extension payload of the audio bitstream, or by writing the updates to a separate metadata track as part of the MP4 file. Without introducing additional system delay, the updates can have a look-ahead time (at the output of the DRC Characteristic A block) equal to the delay of the side chain that generates the DRC gain sequence. A longer look-ahead time improves the initial integrated loudness update at the beginning of the sound program, i.e., it can approach the program loudness of the sound program.
[0030] In a first case, which can be illustrated by FIG. 7, the input audio is live audio simultaneously streamed to the decoder side via a bitstream (e.g., via the Internet). In this case, program loudness cannot be provided during streaming (because the live audio event has not yet ended). In this case, the DRC (applied at the decoder side) is dynamically adjusted, i.e., loudness normalized, based on an in-stream integrated loudness update value, which is a dynamically changing normalization gain that may be equal to the difference between the DRC input loudness target value and the dynamically changing integrated loudness update value, as shown. To limit the rate of change of the integrated loudness update value, the update value sequence may be smoothed at the beginning of the stream rather than the end of the stream. Also, the initial update value (at the beginning of the stream) may take into account the expected loudness of the input audio. For example, the expected loudness may be the result of a carefully performed professional studio setup and pilot measurements of an earlier portion of the input audio that has already elapsed.
[0031] In the second case, the input audio (encoder side) is a live audio recording (rather than live streaming) that is written to an audio file on the encoder side as shown in FIG. 4. In that case, the final integrated loudness update value (the true integrated loudness of the sound program or program loudness) can be written to the file at the end of the recording without the need to rewrite the file. If MPEG-D DRC compliance is desired, this can be achieved by writing the final integrated loudness update value (on the encoder side) into a loudness "box" or field at the ISO Base Media File Format level. This Audio Stream Loudness box type is called ludt. Further referring to FIG. 7, once the encoded audio and its associated encoder-side DRC gain sequence and integrated loudness update value are obtained by the decoder side, the decoder-side processing can apply DRC by determining a DRC gain sequence (using DRC characteristic B) based on a loudness normalized version of the decoded audio signal. This normalization is achieved in this example by adjusting the restored, smoothed instantaneous loudness at the output of inverse characteristic A, preferably using the final integrated loudness update value written into the loudness box. Even if the recording is completed without adding a loudness box to the stream on the encoder side, loudness normalization can still be applied on the decoder side by using the integrated loudness update value in the stream.
[0032] Since the integrated loudness update value within a stream may change slowly over time, for example every 1-10 seconds, normalization effectively shifts the DRC characteristic B accordingly. If the integrated loudness update is based on a short period (elapsed time interval) of the sound program, this shift may be audible at the beginning of the recording or stream during playback of the decoded audio. To limit the rate of change of the integrated loudness update value, the update value itself may be smoothed at the beginning of the recording or streaming rather than at the end.
[0033] In the encoder-side process shown in Figure 6, where the input audio is recorded live to a file, the input audio may be compressed (DRC) at the encoder side using side-chain loudness normalization, and then encoded and written to a file. This process essentially results in a compressed audio output that is comparable, if not essentially the same, to the compressed audio output that would result from the decoder-side process according to Figure 7, where the decoded audio is compressed at the decoder side by loudness normalization based on the integrated loudness update value included in the bitstream. However, deferring loudness normalization to the decoder side as shown in Figure 7 has the advantage that when the recording or event is finished, simply adding the final integrated loudness update value to the MP4 level at the ISO Base Media File Format level will improve the listening experience when the file is played back.
[0034] Reference is now made to Figure 8, which is a flow diagram of a new encoder-side process that can generate both backward-compatible and non-backward-compatible MPEG-D DRC bitstream extensions for decoder-side DRC. A backward-compatible bitstream extension field or payload is one that can be processed by a conventional decoder (decoder-side processing) to perform DRC according to this extension, but without loudness normalization (when applying DRC to the decoded audio signal). An example of such a conventional decoder can be seen in Figure 6. A non-backward-compatible bitstream extension is one that cannot be processed by a conventional decoder (to generate compressed audio). This dual functionality may be enabled as follows:
[0035] A flag may be defined in the bitstream, and may be called, for example, characteristicV1Override. The encoder side can set or clear this flag as follows: To generate a backward-compatible bitstream, the flag is given a first value, such as characteristicV1Override=1, in which case the bitstream also includes a loudness normalization gain, also referred to as encDrcNormGainDb. In this mode, the encoder side processing determines the first DRC gain sequence by applying the first DRC characteristic to the audio signal with loudness normalization using the loudness normalization gain (also referred to herein as encoder-side DRC normalization gain). Referring to Figures 10A and 10B, these are block diagrams of an MPEG-D DRC-compliant audio codec system in which the backward-compatible encoder side generates a backward-compatible bitstream that can be processed by both the new decoder and the legacy decoder. In cases where the input audio is a live recording, an integrated loudness update value is also calculated and provided to the encoder (embedded in the bitstream). The loudness normalization gain may be calculated by subtracting the predicted program loudness value from the DRC input loudness target (e.g., assuming units of dBA), as shown in FIG. 10A.
[0036] The loudness normalization gain encDrcNormGainDb is a value applied in the new backward-compatible encoder-side processing to generate a backward-compatible bitstream, which results in a DRC gain sequence with loudness normalization (for DRC characteristic A). This bitstream can be processed by both the new decoder and a conventional decoder, as shown in FIG. 10B, for example. When this bitstream is processed by a conventional decoder, the decoder does not apply loudness normalization during DRC. When the bitstream is processed by a new decoder that applies loudness normalization during DRC, encDrcNormGainDb is used to cancel, neutralize, or disable the application of encDrcNormGainDb by the backward-compatible encoder in order to apply more accurate loudness normalization using the integrated loudness update value. In other words, the processor of the new decoder cancels the encoder-side DRC normalization gain when applying decoder-side DRC loudness normalization.
[0037] Returning to Figure 8, to enable processing by both legacy and new decoders, the backward compatible bitstream may also include, when the flag has a first value, e.g., characteristicV1Override=1, a first DRC setting field, e.g., UNIDRCCONFEXT_V1, and a second DRC setting field, e.g., UNIDRCCONFEXT_V2. The first DRC setting field instructs the decoder-side process to apply DRC without loudness normalization to the decoded audio signal, e.g., as shown in the legacy decoder block of Figure 10B. The second DRC setting field instructs the decoder-side process to apply DRC with loudness normalization to the decoded audio signal, e.g., as shown in the new decoder block of Figure 10B.
[0038] Continuing with FIG. 8, the new encoder side can create a non-backward compatible MPEG-D DRC bitstream extension (one that cannot be processed by legacy decoders to generate compressed audio) as follows. Note that the encoder side may want to create this DRC bitstream extension if it knows that the bitstream will only be processed by the new decoder side. In such a bitstream, the flag has a second value, e.g., characteristicV1Override=0, and the bitstream does not include the loudness normalization gain (intended for use by the decoder side). In addition, the first DRC setting field, e.g., UNIDRCCONFEXT_V1, is also omitted from the bitstream. FIG. 9 shows a new decoder that can process such a bitstream. In other words, when the flag has a second value, e.g., characteristicV1Override=0, the bitstream includes the second DRC setting field but does not include the first DRC setting field.
[0039] 9 is a flow diagram of a new decoder-side process that can generate a DRC gain sequence using either backward-compatible or non-backward-compatible MPEG-D DRC bitstream extensions. The process may begin by parsing the bitstream to detect a second DRC configuration field, e.g., UNIDRCCONFEXT_V2, and a flag characteristicV1Override. In response to the flag having a first value, e.g., characteristicV1Override=1, the process applies a DRC to the audio signal using DRC characteristic B, e.g., as shown in FIG. 10B (New Decoder Block), with loudness normalization, which uses i) a loudness normalization gain (e.g., encDrcNormGainDb) and ii) multiple instances of an integrated loudness update value, both of which are decoded from the bitstream by a DRC decoder along the audio signal.
[0040] In one aspect, and still referring to FIG. 9 , when the flag has a first value, e.g., characteristicV1Override=1, the index of a first DRC characteristic that may be included in the first DRC settings field is overridden by the index of the first DRC characteristic that is included in the second DRC settings field. For example, an MPEG-D DRC may define DRC characteristics 1 through 6 (also referred to herein as legacy index values or legacy ranges) that are recognizable by a legacy MPEG-D DRC decoder. In this disclosure, in accordance with the extended MPEG-D DRC procedure, those same characteristics are replicated with different index values (also referred to herein as new index values or new ranges), such as 65 through 70. In other words, a legacy characteristic can be referenced either by legacy indexes 1 through 6 or by new indexes 65 through 70, and the parameters of the characteristic remain the same, as shown in the following table: [Table 1]
[0041] When the new encoder-side process generates a backward-compatible bitstream (right side of the flow diagram in FIG. 8 , characteristicV1Override=1), it generates both a first (V1) and a second (V2) DRC setting extension field, where the first DRC setting field points to one or more of the legacy indexes 1-6 instead of any of the new indexes 65-70 to enable backward compatibility with legacy decoders. The V2 extension field may point to one or more of the new index values or one or more of the legacy index values. The new index values effectively inform new decoders (those conforming to the extended MPEG-D DRC procedure of this disclosure) that loudness normalization may be required when generating the second DRC gain sequence. The UNIDRCCONFEXT_V2 extension only corresponds to DRC characteristic indexes 65-70 that require loudness normalization in the decoder.
[0042] The new decoder-side process may decode both the V1 and V2 extension fields, as shown on the right side of Figure 9, and may consequently extract two indices (two different index values) that point to the same DRC characteristic A. In this case, the V2 index is said to override V1, because characteristicV1Override=1, and in that case the new decoder replaces the DRC characteristic index obtained from the UNIDRCCONFEXT_V1 extension with the one obtained from the UNIDRCCONFEXT_V2 extension.
[0043] Returning to Figure 8, when a non-backward compatible bitstream (provided to a new decoder, not a legacy decoder) is generated, the flag characteristicV1Override is set to zero, and a UNIDRCCONFEXT_V2 extension is generated in the bitstream. The UNIDRCCONFEXT_V2 extension contains essentially the same bitstream fields as the UNIDRCCONFEXT_V1 extension. UNIDRCCONFEXT_V1 does not correspond to characteristics 65-70, but the transmitted UNIDRCCONFEXT_V2 does. Since loudness normalization for generating the DRC sequence at the encoder side is not applied in this case (see Figure 7), it is not cancelled out in the decoder (also see Figure 7). This situation is equivalent to setting the normalization gain, e.g., encDrcNormGainDb, to 0 in the decoder-side processing of Figure 10B. When such a bitstream is analyzed by the new decoder-side process, in response to i) the flag having a second value and ii) the index being a first value (e.g., in the range of 65 to 70), the decoder-side process applies a DRC to the audio signal using a second DRC characteristic B and with loudness normalization, which uses the integrated loudness update value but does not use the loudness normalization gain (e.g., the value of encDrcNormGainDb of the additive block is set to zero). In other words, when generating a normalized loudness sequence at the input of DRC characteristic B, encDrcNormGainDb is set to zero.
[0044] However, if the new decoder encounters i) the flag having a second value and ii) the index having a second value different from the first value (e.g., in the range 1-6), the decoder-side processing applies DRC to the audio signal (using the second DRC characteristic B), but without loudness normalization. In other words, referring to Figure 10B, the reconstructed, smoothed instantaneous loudness sequence at the output of inverse characteristic A is not adjusted (before being input to DRC characteristic B). Therefore, the additive block shown in that figure does not exist.
[0045] The following appendix contains a preliminary specification of a proposed method for deferred loudness normalization within the framework of the MPEG-D DRC standard. This document contains an efficient method for generating a bitstream with new information that can be decoded by conventional decoders.
[0046] While particular embodiments have been described and illustrated in the accompanying drawings, it is to be understood that such embodiments are merely exemplary of the broad invention and are not limiting thereof, and that since various other modifications may occur to those skilled in the art, the invention is not limited to the specific constructions and arrangements shown and described. Accordingly, the specification is to be regarded in an illustrative rather than a restrictive sense.
Claims
1. 1. A digital audio method of a decoder, comprising: The method includes obtaining a bitstream; The bitstream comprises: an encoded version of the audio signal; a first dynamic range control (DRC) gain sequence determined by an encoder-side process that applies a first dynamic range control (DRC) characteristic to the audio signal; a loudness normalization gain applied by the encoder when determining the first DRC gain sequence; and an index of the first DRC characteristic, the index identifying or indicating the first DRC characteristic; an integrated loudness value or an integrated loudness update value; The method further comprises:
1. A digital audio method comprising: applying DRC to the audio signal by: i) performing loudness normalization in response to the index having a first value; and ii) not performing loudness normalization in response to the index having a second value.
2. 2. The digital audio method of claim 1, wherein loudness normalization is performed after applying an inverse DRC characteristic to the first DRC gain sequence by using the loudness normalization gain in the bitstream to compensate for or not compensate for the loudness normalization gain applied by the encoder side when determining the first DRC gain sequence.
3. Recovering a loudness sequence by applying the first DRC gain sequence to an inverse of a first DRC characteristic; performing loudness normalization on the reconstructed loudness sequence; generating a second DRC gain sequence by applying the restored loudness sequence to a second DRC characteristic; 2. The digital audio method of claim 1, wherein applying a DRC to the audio signal comprises applying the second DRC gain sequence to the audio signal.
4. 4. The digital audio method of claim 3, wherein the loudness normalization gain is in dB, and performing loudness normalization comprises combining the loudness normalization gain with the reconstructed loudness sequence and the integrated loudness value or integrated loudness update value.
5. 4. The digital audio method of claim 3, wherein performing loudness normalization includes shifting the second DRC characteristic along its input axis by an amount based on the loudness normalization gain and the integrated loudness value or integrated loudness update value.
6. 4. The digital audio method of claim 3, wherein performing the loudness normalization and applying a DRC to the audio signal includes calculating a normalization gain as a difference between a DRC input loudness target and an integrated loudness value or an integrated loudness update value, and adding the normalization gain to the restored loudness sequence to generate the normalized loudness sequence, before applying the normalized loudness sequence to a second DRC characteristic to generate the second DRC gain sequence.
7. extracting an index for the first DRC characteristic from the bitstream and obtaining an inverse of the first DRC characteristic using the extracted index; Recovering a loudness sequence by applying the first DRC gain sequence to the inverse of the first DRC characteristic; If the index has the first value, calculating a normalization gain as the difference between i) the DRC input loudness target and ii) the sum of the integrated loudness value or the integrated loudness update value and the loudness normalization gain used by the encoder-side process, and adding the normalization gain to the restored loudness sequence to generate a normalized loudness sequence; generating a second DRC gain sequence by applying the normalized loudness sequence to a second DRC characteristic; 2. The digital audio method of claim 1, wherein applying a DRC to the audio signal comprises applying the second DRC gain sequence to the audio signal.
8. extracting an index for the first DRC characteristic from the bitstream and obtaining an inverse of the first DRC characteristic using the extracted index; generating a restored loudness sequence by applying the first DRC gain sequence to the inverse of the first DRC characteristic; If the index has the second value, generating a second DRC gain sequence by applying the restored loudness sequence to a second DRC characteristic without performing loudness normalization; 2. The digital audio method of claim 1, wherein applying a DRC to the audio signal comprises applying the second DRC gain sequence to the audio signal.
9. 1. A digital audio method of a decoder, comprising: The method includes obtaining a bitstream; The bitstream comprises: an encoded version of the audio signal; a first dynamic range control (DRC) gain sequence determined by an encoder-side process that applies a first DRC characteristic to the audio signal; an index of the first DRC characteristic, the index identifying or indicating the first DRC characteristic; an integrated loudness value or an integrated loudness update value; a flag, wherein when the flag has a first value, the bitstream includes an encoder-side loudness normalization gain, and when the flag has a second value, the bitstream does not include an encoder-side loudness normalization gain; The digital audio method further comprises applying a DRC to the audio signal according to the flag.
10. applying a DRC to the audio signal in response to the flag having the first value, 10. The digital audio method of claim 9, comprising using a second DRC characteristic with loudness normalization based on i) the encoder-side loudness normalization gain, and ii) the integrated loudness value or the integrated loudness update value.
11. applying a DRC to the audio signal i) in response to the flag having the second value, and ii) when the index has a first value, 10. The digital audio method of claim 9, comprising using a second DRC characteristic with loudness normalization based on the integrated loudness value or the integrated loudness update value but without using an encoder-side loudness normalization gain.
12. i) in response to a flag having a second value, and ii) when the index has a second value, applying a DRC to the audio signal comprises:
10. The digital audio method of claim 9, including using a second DRC characteristic without using loudness normalization.
13. The bitstream includes an encoder-side DRC loudness normalization gain, and applying a DRC to the audio signal comprises:
10. The digital audio method of claim 9, comprising compensating the encoder-side DRC loudness normalization gain by applying a decoder-side DRC loudness normalization gain that is based on the encoder-side DRC loudness normalization gain in the bitstream.
Citation Information
Patent Citations
Efficient drc profile transmission
JP2017534903A