Delayed loudness adjustment for dynamic range control

By deferring loudness adjustment to the decoder side and using integrated loudness measurements, the challenges of improper dynamic range control in live audio streaming and recording are addressed, resulting in improved audio quality and reduced undesirable shifts.

JP2026086552APending Publication Date: 2026-05-26APPLE INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2026-02-04
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Metadata-based dynamic range control (DRC) in audio streaming and recording applications faces challenges when the actual program loudness deviates significantly from predictions, leading to improper adjustment and undesirable loudness shifts, particularly in live streaming and recording scenarios.

Method used

Deferred loudness adjustment is implemented on the decoder side, using integrated loudness measurements and normalization techniques to align DRC gain with actual program loudness, reducing the need for precise upfront predictions and minimizing undesirable shifts.

Benefits of technology

This approach improves audio quality by accurately adjusting dynamic range control based on real-time or post-recording loudness measurements, reducing pumping effects and enhancing the listening experience in live streaming and recording applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086552000001_ABST
    Figure 2026086552000001_ABST
Patent Text Reader

Abstract

It provides a delayed loudness adjustment technique for dynamic range control. [Solution] A bitstream containing an encoded version of the audio signal and an instantaneous loudness sequence of the audio signal is acquired by the decoder. The instantaneous loudness sequence is not loudness normalized. A Dynamic Range Control (DRC) gain sequence is generated by applying the instantaneous loudness sequence with loudness normalization to the DRC characteristics. This DRC gain sequence is applied to the decoded audio signal. Other embodiments are also described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio decoder device, and more particularly to delayed loudness adjustment for dynamic range control.

Background Art

[0002] Sound programs such as music, podcasts, short video clips of live recordings, or feature films have their dynamics (changes in loudness and softness) and a loud part and a soft part that define the dynamic range. In many situations such as listening through headphones in a noisy environment or listening through home loudspeakers at night, it is desirable to reduce the dynamics and dynamic range of the reproduced sound in order to improve the listener's experience. For that purpose, a dynamic range compressor is used. This compressor is a digital signal processor that applies a time-varying gain to an input that is a digital audio signal (of a sound program), amplifying the soft part of the audio signal and attenuating the loud part. In order to avoid audible pumping artifacts that may result from compression of the dynamic range of the audio signal, a loudness normalization process can be performed to "match" the input audio signal to the compression characteristic or profile while compressing the audio signal according to the compression characteristic. This process can be performed by offsetting the instantaneous loudness of the input audio signal by the program loudness of that signal, where the program loudness is a calculated value (also referred to as integrated loudness) intended to represent the overall loudness of the sound program.

Summary of the Invention

[0003] Audio coding standards define methods of dynamic range compression in which dynamic range control (DRC) gain is generated on the encoder side when the sound program is created or prepared for distribution or storage / archiving. In this specification, DRC gain refers to a DRC gain sequence that is temporally aligned with the relevant sound program, such that one or more gain values ​​in the sequence are applied to corresponding digital audio frames in the sound program. The DRC gain sequence is then formatted into one or more bitstreams, for example, as metadata associated with the sound program. The decoder side receives the bitstreams and, if desired by the decoder (typically during playback of the decoded audio signal), applies the DRC gain in the stream to compress the dynamic range of the decoded audio signal. An advantage of such metadata-based methods is improved quality due to the longer advance time intervals available for offline coding of the DRC gain compared to real-time compression. Another advantage is that the compression characteristics can be controlled on the encoder side, for example, by the expertise of the sound program creator or distributor.

[0004] Metadata-based DRC in online applications (e.g., live audio streaming and recording of live audio to files) presents challenges when the program loudness of the sound program being streamed for playback or written for storage is not yet known (because the sound program has not yet terminated). This is because if the actual program loudness of the sound program (which can only be determined after the sound program has terminated) deviates significantly from what is expected or predicted, the compressor characteristics may not be properly adjusted (or loudness normalized).

[0005] Some aspects of the disclosure herein relate to novel digital signal processing methods that defer loudness adjustment (loudness normalization) of dynamic range control (DRC) from the encoder side to the decoder side. Other aspects relate to techniques for modifying compressor characteristics on the decoder side when metadata-based DRC gain sequencing processing is used for loudness normalization. These aspects are particularly useful for applications such as live streaming and live recording to files.

[0006] The above summary does not provide an exhaustive list of all embodiments of this disclosure. This disclosure is intended to include all implementable systems and methods from all preferred combinations of the various embodiments summarized above, as well as those disclosed in the following “Modes for Carrying Out the Invention” and those specifically indicated in the “Claims” section. Such combinations may have specific advantages not specifically listed in the above summary.

[0007] Some aspects of this disclosure are described for illustrative purposes only and are not intended to limit the descriptions to the figures in the accompanying drawings where similar reference numerals indicate similar elements. It should be noted that references to "an" or "one" aspects in this disclosure do not necessarily refer to the same aspect, but rather to at least one. Furthermore, for the sake of brevity and to reduce the total number of figures, a given figure may be used to illustrate features of multiple aspects of this disclosure, and not all elements in the figure are required for the given aspect. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an exemplary DRC characteristic curve. [Figure 2] This is a block diagram of an audio codec system where DRC is applied on the decoder side, and loudness normalization is not performed on the encoder side. [Figure 3]This is a block diagram of an audio codec system suitable for live streaming, where DRC is applied on the decoder side, and loudness normalization is not performed on the encoder side. [Figure 4] This is a block diagram of an audio codec system suitable for live recording to storage or an archive, where DRC is applied on the decoder side and loudness normalization is not performed on the encoder side. [Figure 5] This diagram shows a part of an MPEG-D DRC-compliant audio codec system where DRC is applied on the decoder side. [Figure 6] This figure shows a part of an MPEG-D DRC-compliant audio codec system, where loudness normalization is provided on the encoder side and DRC is applied on the decoder side. [Figure 7] This shows a part of an MPEG-D DRC-compliant audio codec system that applies DRC with loudness normalization on the decoder side. [Figure 8] This is a flowchart of the new encoder-side processing that can generate backward-compatible and backward-incompatible MPEG-D DRC bitstream extensions. [Figure 9] This is a flowchart of a new decoder-side processing that can generate a DRC gain sequence using either a backward-compatible or non-backward-compatible MPEG-D DRC bitstream extension. [Figure 10A] This is a block diagram of an MPEG-D DRC-compliant audio codec system in which the backward-compatible encoder generates a backward-compatible bitstream that is processed by both the new and conventional decoders. [Figure 10B] This is a block diagram of an MPEG-D DRC-compliant audio codec system in which the backward-compatible encoder generates a backward-compatible bitstream that is processed by both the new and conventional decoders. [Modes for carrying out the invention]

[0009] Several aspects of this disclosure are described herein with reference to the accompanying drawings. Wherever the shape, relative position, and other aspects of the described parts are not expressly specified, the scope of the invention is not limited to the illustrated parts, and they are for illustrative purposes only. Furthermore, while numerous details are described, it will be understood that some aspects of this disclosure can be carried out without these details. In other examples, well-known circuits, structures, and techniques are not shown in detail so as not to hinder the understanding of this specification.

[0010] To properly apply dynamic range control to an audio signal, the compressor characteristics (DRC characteristics, DRC profile) should be "matched" to the loudness level range of the audio signal. For example, referring to Figure 1, the matching is done along the input level axis such that the zero crossing of the DRC characteristic curve is approximately in the center of the audio signal's loudness level range. The level at the zero crossing is also referred to as the DRC input loudness target, and in the exemplary set of characteristic curves shown in Figure 1, this level is approximately -31 dB. The center of the loudness level range may be, for example, the average level of a sound program, or the average dialogue level within a sound program. In this specification, the process for achieving such matching is referred to as loudness normalization of the audio signal to a given loudness target, related to the DRC. For example, the loudness of an audio signal (sound program) may be a single value known as integrated loudness. Integrated loudness is a measure of the loudness of an audio signal, similar to the root mean square (RMS), but more faithful from the perspective of human hearing. Integrated loudness can be equivalent to program loudness in that it measures how loud a sound program is over its entire duration. To achieve loudness normalization, if the integrated loudness is given in decibels (dB), it can be subtracted from the DRC input loudness target to derive the normalization gain in dB. This normalization gain is added to the output of a loudness model that calculates the instantaneous loudness of the audio signal (sound program). The instantaneous loudness may be a sequence of loudness values ​​calculated based on each digital audio frame that makes up the input digital audio signal (and representing human-perceived loudness). Another way to achieve loudness normalization is to shift the DRC characteristic curve shown in Figure 1 to the right or left (by the amount of the normalization gain).In the example in Figure 1, the curve is shifted to the left by -31 dB (the loudness target in this example), and is therefore well-suited (and thus directly applicable) to sound programs with an integrated loudness of -31 dBA (A-weighted) or LKFS (loudness K-weighted level full scale). In other words, the normalization gain in that case is zero dBA.

[0011] If the integrated loudness of an audio program is still unknown during dynamic range control signal processing, a prediction must be made to apply loudness normalization, as is the case with live audio. However, predictions can be inaccurate, resulting in a DRC gain that contains undesirable biases, or a DRC gain that generates a pumping effect, which is an undesirable loudness shift between the uncompressed and compressed portions of the audio signal.

[0012] To reduce the possibility of undesirable loudness shifts, one aspect of the disclosure herein applies loudness normalization to the DRC on the decoder side rather than the encoder side of the audio codec system or method. An example of an audio codec system and associated method is shown in the hardware block diagram of Figure 2. Various hardware blocks of the audio codec system and method may be implemented by a programmed processor. In such a method, integrated loudness (necessary for loudness normalization performed in relation to the DRC for playback or archiving / storage of the decoded audio signal) can be obtained in at least two examples described below in relation to Figures 3 and 4.

[0013] First, looking at Figure 2, the audio codec system has an encoder side, which may be implemented, for example, in one or more servers, by one or more processors that execute or are configured by instructions stored in memory, collectively referred to herein as “programmed processors”. The upper audio signal processing path includes a side chain, which includes a loudness model that calculates or estimates the instantaneous loudness of a digital input audio signal (sequence), also referred herein as a sound program. This estimation is based on a perceived loudness scale (such as a Thorne scale) and is therefore approximately logarithmic. To smooth the instantaneous loudness sequence over time, a smoothing filter can be applied as shown in the figure. As a result, regions of the input audio sequence where changes in compression gain are not desired are smoothed, but macrodynamic loudness transitions remain unaffected.

[0014] The lower audio signal processing path includes a delay block that delays the input audio sequence to compensate for delays caused by the side chains. The smoothed loudness sequence and the delayed input audio sequence are then fed to the encoder.

[0015] The encoder may perform a bitrate reduction operation on one or both inputs to generate one or more bitstreams containing bitrate-reduced versions of one or both inputs. These one or more bitstreams may then be transmitted to the decoder (e.g., over the Internet) or written to a file for storage or archiving until accessed for decoder-side processing. Smoothed loudness (referring to a sequence of smoothed loudness values, or a single smoothed loudness value) may be carried as metadata within the same bitstream as the delayed input audio sequence, for example, associated with the "corresponding" Advanced Audio Coding (AAC) audio frame. This is also referred to as existing within the audio layer. Alternatively, the smoothed loudness sequence and other loudness values, such as integrated loudness update values ​​and DRC payloads (as will be discussed further below), may be transmitted at a higher layer, such as the file format level, rather than within the audio layer. In either case, one or more bitstreams are generated, in which the encoded audio is provided along with associated metadata, such as the smoothed loudness sequence, or, in other embodiments, along with instructions to the decoder for applying the DRC gain sequence supplied by the encoder, as described below.

[0016] The decoder side may also be implemented as a programmed processor, such as one or more processors executed or configured by instructions stored in memory as part of an audio playback device. Note that the decoder side processing may be implemented within the same audio playback device as the encoder side processing. Alternatively, the decoder side processing may be implemented within a separate audio playback device from the programmed processor that performs the encoder side processing. Examples of audio playback devices include smartphones, tablet computers, digital media players, headsets, or vehicle infotainment systems. On the decoder side, in order to restore the smoothed loudness sequence and the delayed input audio sequence, the decoder deactivates the encoder's bitrate reduction operation. Once the smoothed loudness value is decoded, it is then mapped to the "corresponding" DRC, or compression gain value. This mapping is a memoryless input / output function that implements, for example, one of the curves shown in Figure 1 (or any other desired curve). This mapping constitutes a compressor characteristic or compressor profile (DRC characteristic), the output of which is a time-varying gain (a sequence of DRC gain values) that is a function of the time-varying input loudness level. This mapping may also include a conversion from a logarithmic loudness domain to a linear domain of DRC gain. Then, if compression is desired, the DRC gain values ​​(sequence) are applied to the decoded audio signal as indicated by the multiplication symbols in the figure. Although not shown, the compressed audio is then passed to a playback processing block, which may ultimately generate a transducer (speaker) driver input signal that converts the compressed audio into sound.

[0017] As can be seen in Figure 2, the smoothed loudness sequence is adjusted, or normalized, on the decoder side before being input to the DRC mapping block. For example, the normalization gain in dB units can be derived by subtracting the invariant integrated loudness (single value) from the DRC input loudness target (in the dB range). This normalization gain is added to each of the smoothed loudness values ​​in the smoothed loudness sequence to generate the normalized loudness sequence used in DRC processing. Such DRC processing has at least two applications, e.g., live or real-time streaming, and live recording to a file for storage or archiving.

[0018] In one such application, referring to Figure 3, the input audio on the encoder side is, for example, a live or real-time digital audio recording streamed to the decoder side over the internet. The input audio includes audio capture of the live or real-time event, which occurs simultaneously with encoding and bitstream transmission. Therefore, a single integrated loudness value representing the entire sound program cannot be calculated until the live event has finished. In the meantime, an integrated loudness measurement block within the encoder collects samples of the live audio, which are sent to the encoder with a time delay for time matching, over time intervals longer than a single audio frame, e.g., several seconds, and calculates a loudness measurement for that time interval. This block then "integrates" or collects some of these measurements, working backward from the beginning of the sound program, and calculates an integrated loudness update value, for example, by averaging them. The integrated loudness update value may be a measurement of integrated loudness only for the portion of the sound program that has been played or streamed up to the current update. For example, periodically, this measurement is repeated to generate a "moving average" integrated loudness and sends the latest integrated loudness update value (which is a single value) to the decoder side. When used herein, the term “moving average” should be noted as requiring only a few measurements of the loudness of the sound program from the beginning of the program to the current update, based on the collection of loudness measurements, including a statistical evaluation of the collected loudness measurements, and not on performing an actual average. The update value (moving average) may be calculated and then provided as multiple instances of an integrated loudness update value field, as part of a bitstream that also includes the encoded sound program (encoded audio signal), such that adjacent instances in the bitstream are separated by only 1 to 10 seconds over the duration of the sound program.

[0019] Note that the term "integrated loudness update value" may also be referred to as moving average loudness or "partial integrated loudness". At the end of a sound program, the last or final loudness update value may represent the loudness of the entire sound program (e.g., integrated loudness or program loudness, also referred to as such, as described in the ITU-R BS.1770-4 (10 / 2015) algorithm for measuring audio program loudness and true peak audio level).

[0020] On the decoder side, the decoder obtains the bitstream, extracts the integrated loudness update value therefrom, and then the decoder-side processing applies it to perform loudness normalization on the DRC processing. This may be done, for example, by adding a single loudness normalization gain value (e.g., the difference between the DRC input loudness target and the integrated loudness update value) to the decoded or restored instantaneous loudness sequence and then inputting it to the DRC characteristic mapping block. Alternatively, loudness normalization may be done by shifting the DRC characteristic along the input axis by an amount equal to the loudness normalization gain value. The loudness normalization gain may be updated periodically during the transmission of the bitstream (sound program) using the latest partial integrated loudness value (integrated loudness update value) calculated on the encoder side for the elapsed portion of the live event.

[0021] In another application example, referring to FIG. 4, the input audio on the encoder side is a live or real-time digital audio recording of an event, which is written to a file for archival or storage purposes at the end of the recording (when the event ends). A single integrated loudness value representing the program loudness of the entire live audio event can be calculated by the integrated loudness model block at the end of the recording and provided to the encoder as soon as the event ends. The encoder writes the integrated loudness value to the file along with the encoded version of the live audio and the encoded version of the instantaneous (and smoothed) loudness sequence calculated by the loudness model (based on the same live audio). On the decoder side, the decoder obtains the file (bitstream), decodes the input audio and the instantaneous loudness sequence from the file, and extracts the integrated loudness value from the file. The decoder-side processing then performs loudness normalization on the decoded instantaneous loudness sequence using the integrated loudness value, inputs it to the DRC (compression) mapping block, and then, during playback (if compression is desired), the output of this block is applied to the decoded input audio.

[0022] In one aspect, the smoothing filter is a non-linear filter as described in U.S. Patent No. 10,109,288. A useful property of this filter is that it can level-shift its output by the same amount as the input. That is, if f(x) is defined as a non-linear function, x(n) as the input signal, and y(n) as the output, it can be described as follows. y(n)=f(x(n))

[0023] When the shift of the input signal is given by ΔL and the output is shifted by ΔL, f(x) satisfies the shift property, and this can be expressed mathematically as follows. y(n)+ΔL=f(x(n)+ΔL)

[0024] This is beneficial because it completely avoids the side-chain processing on the encoder side that has a dependency on the absolute loudness value.

[0025] Another aspect of the disclosure herein is a method for applying DRC in accordance with the MPEG-D DRC standard ISO / IEC, “Information technology - MPEG Audio Technologies - Part 4: Dynamic Range Control”, ISO / IEC 23003-4:2020 (“MPEG-D DRC”), extended to accommodate loudness normalization on the encoder side. Figure 5 shows a simplified block diagram of part of the MPEG-D DRC processing, where the DRC gain is generated and applied based on decoding the DRC gain from metadata in the bitstream obtained from the encoder side. MPEG-D DRC provides a default DRC characteristic and a flexible method for encoding parameterized characteristics.

[0026] In Figure 5, the encoder applies the smoothed instantaneous loudness sequence (calculated for the input audio sequence) to the selected DRC characteristic (also referred to as the "mapping block" used above in relation to Figure 2). The output of the DRC characteristic mapping block generates a DRC gain sequence, which is supplied to the DRC encoder. The DRC encoder performs bitrate reduction to encode the input sequence into one or more bitstreams, which are then transmitted or otherwise made available to the decoder. At the decoder side, the DRC decoder decrypts the bitrate reduction encoding to restore the DRC gain sequence (decoded DRC gain sequence). The decoded DRC gain sequence is then applied to the decoded audio signal (if compression is desired).

[0027] MPEG-D DRC also supports a type of decoder-side processing that changes the DRC characteristics applied to compress the sound program from those used on the encoder side (to calculate the DRC gain sequence inserted into the bitstream as metadata), as shown in Figure 5, to others that may be selected by decoder-side processing based on the current playback or listening conditions. To achieve this, first, the encoder-side DRC gain sequence is applied to inverse characteristic A, as shown in Figure 6. Inverse characteristic A is the reciprocal of DRC characteristic A, which is applied on the encoder side to generate the encoder-side DRC gain sequence. An index (identifier or pointer) to DRC characteristic A (the one used on the encoder side to generate the DRC gain sequence) may be provided in the bitstream so that the decoder side can identify inverse characteristic A. When the DRC gain sequence is applied as input to inverse characteristic A, the result is a smoothed instantaneous loudness sequence. Ignoring quantization effects, the restored loudness sequence (at the output of the inverse characteristic A block) is essentially the smoothed loudness sequence used by the encoder-side processing. As a result, the restored loudness sequence can be applied to a second DRC characteristic B to generate a second DRC gain sequence that may be more suitable (than DRC characteristic A) for compressing the decoded audio signal. Then, (for example, if compression is desired during playback) the second DRC gain sequence is applied to the decoded audio.

[0028] According to one aspect of the disclosure herein, the loudness normalization of the encoder-side chain shown in Figure 6 is replaced using the method shown in Figure 2. That is, an offset (normalization gain) based on integrated loudness is applied on the decoder side rather than the encoder side. Figure 7 shows a block diagram of such a system. In this specification, this system is also referred to as an extended MPEG-D DRC compliant system (hereinafter also referred to as having a "new" encoder and a "new" decoder). Such a system has a block on the encoder side called an integrated loudness measurement, the output of which provides an integrated loudness update value as discussed above with respect to Figure 3. This integrated loudness update value is provided to an audio encoder. In this specification, this encoder is a DRC encoder that also encodes a DRC gain sequence (in addition to the input audio). The DRC gain sequence may be determined as discussed above with respect to Figure 6. The encoded DRC gain sequence and integrated loudness update value are provided to the decoder side via one or more bitstreams. The DRC gain sequence may be formatted as metadata associated with the encoded input audio, which is also provided to the decoder side.

[0029] Integrated loudness measurement is a moving measurement of integrated loudness (also referred to herein as a moving average) that begins acquisition at the start of the sound program and continues over time to "integrate" the audio signal of the sound program for the purpose of calculating integrated loudness only for the portion of the sound program that has progressed. As the audio signal (sound program) continues, the integrated loudness measurement generates updated values, for example, periodically, for example, every 10 seconds. These integrated loudness update values ​​are written to the bitstream (for example, by a DRC encoder). In MPEG-D DRC, this can be done by writing the updated values ​​to an extended field or extended payload of the audio bitstream, or by writing the updated values ​​to a separate metadata track as part of the MP4 file. Without introducing extra system delay, the updates may have a look-ahead time equal to the delay of the side chain that generates the DRC gain sequence (at the output of DRC characteristic A block). A longer look-ahead time improves the initial integrated loudness update at the beginning of the sound program, that is, it may approach the program loudness of the sound program.

[0030] In the first example illustrated by Figure 7, the input audio is live audio simultaneously streamed to the decoder via a bitstream (e.g., over the Internet). In this case, program loudness cannot be provided during streaming (because the live audio event has not yet finished). In this case, DRC (applied on the decoder side) is dynamically adjusted, i.e., loudness normalization is performed, based on an in-stream integrated loudness update value, which is a dynamically changing normalization gain that may be equal to the difference between the DRC input loudness target value and the dynamically changing integrated loudness update value, as illustrated. To limit the rate of change of the integrated loudness update value, the update value sequence may be smoothed at the beginning of the stream rather than the end of the stream. The initial update value (beginning of the stream) may also take into account the expected loudness of the input audio. For example, the expected loudness may be the result of a carefully performed professional studio setup and pilot measurements of the initial portion of the input audio that has already passed.

[0031] In the second example, the input audio (encoder side) is a live audio recording written to an audio file on the encoder side, as shown in Figure 4 (not a live stream). In this case, the final integrated loudness update value (the true integrated loudness or program loudness of the sound program) can be written to the file at the end of the recording without the need to rewrite the file. If the goal is to comply with MPEG-D DRC, this can be achieved (on the encoder side) by writing the final integrated loudness update value to a loudness "box" or field at the ISO-based media file format level. This Audio Stream Loudness box type is called ludt. Referring further to Figure 7, once the encoded audio and its associated encoder-side DRC gain sequence and integrated loudness update value are acquired by the decoder side, the decoder side can apply DRC by determining the DRC gain sequence (using DRC characteristic B) based on the loudness normalized version of the decoded audio signal. In this example, this normalization is achieved by adjusting the restored, smoothed instantaneous loudness at the output of inverse characteristic A, preferably using the final integrated loudness update value written to the loudness box. Even if recording ends without adding a loudness box to the stream on the encoder side, loudness normalization can be applied on the decoder side by using the integrated loudness update value in the stream.

[0032] The integrated loudness update value within a stream may change slowly over time, for example, every 1 to 10 seconds, and accordingly, normalization effectively shifts DRC characteristic B. If the integrated loudness update is based on a short period (elapsed time interval) of the sound program, this shift may be audible at the beginning of the recording or stream during playback of the decoded audio. To limit the rate of change of the integrated loudness update value, the update value itself may be smoothed at the beginning of the recording or streaming rather than at the end.

[0033] In the encoder-side processing shown in Figure 6, where the input audio is a live recording to a file, the input audio may be compressed using side-chain loudness normalization (DRC) on the encoder side, then encoded and written to the file. This process results in a compressed audio output that is comparable, if not essentially the same, to the compressed audio output obtained from the decoder-side processing shown in Figure 7 (where the decoded audio is compressed on the decoder side by loudness normalization based on the integrated loudness update value contained in the bitstream). However, deferring loudness normalization to the decoder side, as shown in Figure 7, has the advantage that when the recording or event ends, the final integrated loudness update value is simply added to the MP4 level, which is the level of the ISO-based media file format, improving the listening experience when the file is played back.

[0034] Refer to Figure 8 here. This is a flowchart of a new encoder-side processing that can generate both backward-compatible and backward-compatible MPEG-D DRC bitstream extensions for DRC by the decoder. A backward-compatible bitstream extension field or payload can be processed by a conventional decoder (decoder-side processing) to perform DRC as per this extension, but without loudness normalization (when applying DRC to the decoded audio signal). An example of such a conventional decoder can be shown in Figure 6. A backward-compatible bitstream extension is one that cannot be processed by a conventional decoder (to generate compressed audio). This dual functionality may be enabled as follows.

[0035] A flag may be defined within the bitstream, for example, called characteristicV1Override. The encoder can set or clear this flag as follows. To generate a backward-compatible bitstream, a first value such as characteristicV1Override=1 is given to the flag, in which case the bitstream also includes loudness normalization gain, also referred to as encDrcNormGainDb. In this mode, the encoder processing determines a first DRC gain sequence by applying the audio signal with loudness normalization to a first DRC characteristic using the loudness normalization gain (also referred to herein as encoder-side DRC normalization gain). Referring to Figures 10A and 10B, these are block diagrams of an MPEG-D DRC-compliant audio codec system in which the backward-compatible encoder generates a backward-compatible bitstream that is processed by both the new and conventional decoders. In cases where the input audio is a live recording, an integrated loudness update value is also calculated and provided to the encoder (incorporated into the bitstream). The loudness normalization gain may be calculated by subtracting the predicted program loudness value from the DRC input loudness target (for example, assuming units in dBA), as shown in Figure 10A.

[0036] The loudness normalization gain encDrcNormGainDb is a value applied in the new backward-compatible encoder-side processing to generate a backward-compatible bitstream, which in turn obtains a DRC gain sequence with loudness normalization (for DRC characteristic A). This bitstream can be processed by both the new and conventional decoders, as shown in Figure 10B, for example. When this bitstream is processed by the conventional decoder, it does not apply loudness normalization during DRC. When the bitstream is processed by the new decoder, which applies loudness normalization during DRC, encDrcNormGainDb is used to cancel, neutralize, or disable the application of encDrcNormGainDb by the backward-compatible encoder in order to apply more accurate loudness normalization using the integrated loudness update value. In other words, the new decoder's processor cancels out the encoder-side DRC normalization gain when applying decoder-side DRC loudness normalization.

[0037] Returning to Figure 8, in order to enable processing by both conventional and new decoders, the backward-compatible bitstream may also include a flag having a first value, e.g., characteristicV1Override=1, a first DRC setting field, e.g., UNIDRCCONFEXT_V1, and a second DRC setting field, e.g., UNIDRCCONFEXT_V2. The first DRC setting field instructs the decoder-side processing to apply DRC to the decoded audio signal without loudness normalization, for example, as shown in the conventional decoder block of Figure 10B. The second DRC setting field instructs the decoder-side processing to apply DRC to the decoded audio signal with loudness normalization, for example, as shown in the new decoder block of Figure 10B.

[0038] Referring further to Figure 8, the new encoder can create a non-backward compatible MPEG-D DRC bitstream extension (which cannot be processed by conventional decoders to generate compressed audio), as shown below. Note that the encoder may wish to create this DRC bitstream extension if it is aware that the bitstream will be processed only by the new decoder. In such a bitstream, the flag has a second value, for example, characteristicV1Override=0, and the bitstream does not include the loudness normalization gain (intended for use by the decoder). In addition, the first DRC setting field, for example, UNIDRCCONFEXT_V1, is also omitted from the bitstream. Figure 9 shows a new decoder capable of processing such a bitstream. In other words, when the flag has a second value, for example, characteristicV1Override=0, the bitstream includes the second DRC setting field but does not include the first DRC setting field.

[0039] Figure 9 is a flowchart of a new decoder-side processing that can generate a DRC gain sequence using either a backward-compatible or non-backward-compatible MPEG-D DRC bitstream extension. Processing may begin by analyzing the bitstream to detect a second DRC configuration field, e.g., UNIDRCCONFEXT_V2, and the flag characteristicV1Override. In response to the flag having a first value, e.g., characteristicV1Override=1, processing applies DRC to the audio signal using DRC characteristic B and with loudness normalization, as shown, for example, in Figure 10B (new decoder block), where this loudness normalization uses i) loudness normalization gain (e.g., encDrcNormGainDb) and ii) a combined loudness update value for multiple instances (both of which are decoded from the bitstream by the DRC decoder along the audio signal).

[0040] In one embodiment, continuing with reference to Figure 9, when the flag has a first value, for example characteristicV1Override=1, the index of a first DRC characteristic that may be included in a first DRC setting field is overridden by the index of a first DRC characteristic included in a second DRC setting field. For example, MPEG-D DRC may define DRC characteristics 1-6 (also referred herein as conventional index values ​​or conventional ranges) that are recognizable by a conventional MPEG-D DRC decoder. In this disclosure, following the Extended MPEG-D DRC procedure, the same characteristics are duplicated with different index values ​​(also referred herein as new index values ​​or new ranges), for example, 65-70. In other words, conventional characteristics can be referenced by either conventional indices 1-6 or new indices 65-70, and the characteristics parameters remain the same, as shown in the table below. [Table 1]

[0041] The new encoder processing, when generating a backward-compatible bitstream (right side of the flowchart in Figure 8, characteristicV1Override=1), generates both a first (V1) and a second (V2) DRC setting extension field. The first DRC setting field points to one or more of the conventional indices 1-6, rather than any of the new indices 65-70, to enable backward compatibility with conventional decoders. The V2 extension field may point to one or more of the new index values, or one or more of the conventional index values. The new index values ​​effectively inform the new decoder (compliant with the extended MPEG-D DRC procedure of this disclosure) that loudness normalization may be required when generating the second DRC gain sequence. Only the UNIDRCCONFEXT_V2 extension corresponds to the DRC characteristic indices 65-70 that require loudness normalization within the decoder.

[0042] The new decoder processing may decode both the V1 and V2 extension fields, as shown on the right side of Figure 9, and as a result, two indices (two different index values) that indicate the same DRC characteristic A may be extracted. In this case, the V2 index is said to override V1. This is because characteristicV1Override=1, in which case the new decoder replaces the DRC characteristic index obtained from the UNIDRCCONFEXT_V1 extension with the one obtained from the UNIDRCCONFEXT_V2 extension.

[0043] Returning to Figure 8, when a non-backward compatible bitstream (provided to the new decoder, not the conventional decoder) is generated, the flag characteristicV1Override is set to zero, and the UNIDRCCONFEXT_V2 extension is generated in the bitstream. The UNIDRCCONFEXT_V2 extension contains substantially the same bitstream fields as the UNIDRCCONFEXT_V1 extension. UNIDRCCONFEXT_V1 does not correspond to characteristics 65-70, but the transmitted UNIDRCCONFEXT_V2 does. Loudness normalization for generating the DRC sequence on the encoder side is not applied in this case (see Figure 7), and therefore is not canceled out in the decoder (see Figure 7 again). This situation is equivalent to setting the normalization gain, e.g., encDrcNormGainDb, to 0 in the decoder-side processing in Figure 10B. When such a bitstream is analyzed by the new decoder processing, in response to i) the flag having a second value and ii) the index having a first value (e.g., in the range of 65 to 70), the decoder processing applies DRC to the audio signal using the second DRC characteristic B and with loudness normalization, which uses the integrated loudness update value but does not use the loudness normalization gain (e.g., the value of encDrcNormGainDb in the additive block is set to zero). In other words, if a normalized loudness sequence is generated at the input of DRC characteristic B, encDrcNormGainDb is set to zero.

[0044] However, if the new decoder encounters i) a flag having a second value, and ii) an index having a second value different from the first value (for example, in the range 1-6), the decoder processing will apply DRC to the audio signal (using the second DRC characteristic B), but without loudness normalization. In other words, referring to Figure 10B, the restored, smoothed instantaneous loudness sequence at the output of inverse characteristic A is not adjusted (before being input to DRC characteristic B). Therefore, the additive block shown in that figure does not exist.

[0045] The following appendix contains a provisional specification for a proposed method for deferred loudness normalization within the framework of the MPEG-D DRC standard. This document includes an efficient method for generating a bitstream using new information that can be decoded even by conventional decoders.

[0046] While specific embodiments have been described and illustrated in the accompanying drawings, these embodiments are merely illustrative of the broader invention and do not limit it. Furthermore, various other modifications can be conceived by those skilled in the art. Therefore, the present invention is not limited to the specific configurations and arrangements illustrated and described. Accordingly, this specification should be considered illustrative, not restrictive.

Claims

1. Processor and A memory that internally stores instructions for configuring the processor to acquire a bitstream. An audio decoder device comprising, wherein the bitstream is The encoded version of the audio signal, The first dynamic range control, i.e., DRC, gain sequence, is determined by encoder-side processing that applies the aforementioned audio signal to a first DRC characteristic. The loudness normalization gain applied by the encoder when determining the first DRC gain sequence, An index of the first DRC characteristic, wherein the index identifies or indicates the first DRC characteristic, Multiple instances of the integrated loudness update value over time, including, Audio decoder device.

2. The audio decoder device according to claim 1, wherein the processor performs loudness normalization when applying DRC to the audio signal, depending on whether the index has a first value.

3. The audio decoder device according to claim 1, wherein the bitstream instructs the processor to perform loudness normalization by applying an inverse DRC characteristic to the DRC gain sequence, and then canceling or revokeing the loudness normalization applied when the encoder determines the DRC gain sequence, using the loudness normalization gain in the bitstream.

4. The memory internally stores and holds instructions, and the instructions cause the processor to The loudness sequence is reconstructed by applying the first DRC gain sequence to the reciprocal of the first DRC characteristic. Loudness normalization is performed on the restored loudness sequence. A second DRC gain sequence is generated by applying the restored loudness sequence to a second DRC characteristic. The second DRC gain sequence is applied to the audio signal. Configure it as follows: The audio decoder device according to any one of claims 1 to 3.

5. The audio decoder device according to claim 4, wherein the loudness normalization gain is in units of dB, and performing loudness normalization includes combining the loudness normalization gain with instances of the restored loudness sequence and the integrated loudness update value.

6. The audio decoder device according to any one of claims 1 to 4, wherein the loudness normalization includes shifting the second DRC characteristic along the input axis by an amount based on the loudness normalization gain and the integrated loudness update value instance.

7. The audio decoder device according to any one of claims 1 to 6, wherein the processor calculates an update to the normalization gain for each instance of the integrated loudness update value as the difference between the DRC input loudness target and the instance of the integrated loudness update value, adds the normalization gain to the restored loudness sequence to generate a normalized loudness sequence, and then applies the normalized loudness sequence to the second DRC characteristic to generate the second DRC gain sequence.

8. The audio decoder device according to any one of claims 1 to 7, wherein adjacent instances of the integrated loudness update value are separated by only 1 to 10 seconds.

9. The audio decoder device according to any one of claims 1 to 8, wherein the integrated loudness update value represents the moving average integrated loudness of the audio signal.

10. The aforementioned processor, Extract the index to the first DRC characteristic from the bitstream, and use the extracted index to obtain the reciprocal of the first DRC characteristic. The loudness sequence is reconstructed by applying the first DRC gain sequence to the reciprocal of the first DRC characteristic. If the index has a first default value, for each instance of the integrated loudness update value, the normalization gain update value is calculated as the difference between i) the DRC input loudness target and ii) the sum of the instance of the integrated loudness update value and the encoder-side loudness normalization gain used by the encoder-side processing, and the normalization gain update value is added to the restored loudness sequence to generate a normalized loudness sequence. A second DRC gain sequence is generated by applying the normalized loudness sequence to a second DRC characteristic. The second DRC gain sequence is applied to the audio signal. The audio decoder device according to claim 1, configured as described above.

11. The audio decoder device according to claim 10, wherein the processor is configured to generate the second DRC gain sequence by applying the restored loudness sequence to the second DRC characteristic without loudness normalization when the index has a second specified value.

12. Processor and A memory that internally stores instructions for configuring the processor to acquire a bitstream, An audio decoder device comprising, wherein the bitstream is The encoded version of the audio signal, The first dynamic range control, i.e., DRC, gain sequence, is determined by encoder-side processing that applies the aforementioned audio signal to a first DRC characteristic. An index of the first DRC characteristic, wherein the index identifies or indicates the first DRC characteristic, Multiple instances of the integrated loudness update value over time, A flag, wherein when the flag has a first value, the bitstream includes encoder-side loudness normalization gain, or when the flag has a second value, the bitstream does not include the encoder-side loudness normalization gain. including, Audio decoder device.

13. The audio decoder device according to claim 12, in response to the flag having the first value, the processor applies DRC to the audio signal using a second DRC characteristic and with loudness normalization, wherein the loudness normalization uses i) the encoder-side loudness normalization gain and ii) the plurality of instances of the integrated loudness update value.

14. i) in response that the flag has the second value, ii) when the index has the first value, the processor applies DRC to the audio signal using a second DRC characteristic and with loudness normalization, wherein the loudness normalization uses the plurality of instances of the integrated loudness update value but does not use the encoder-side loudness normalization gain, according to claim 12.

15. The audio decoder device according to claim 14, wherein, in response to the index being a second value different from the first value, the processor applies DRC to the audio signal using the second DRC characteristic but without loudness normalization.

16. Processor and A memory that internally stores instructions for configuring the processor to acquire a bitstream, An audio decoder device comprising, wherein the bitstream is The encoded version of the audio signal, The first dynamic range control, i.e., DRC, gain sequence, is determined by encoder-side processing that applies the aforementioned audio signal to a first DRC characteristic. An index of the first DRC characteristic, wherein the index identifies or indicates the first DRC characteristic, Multiple instances of the integrated loudness update value over time, A flag is included, and when the flag has a first value, the processor replaces some or all of the conventional DRC characteristic index values ​​of the conventional extended payload in the bitstream with the DRC characteristic index values ​​included in the new extended payload in the bitstream. Audio decoder device.

17. Processor and A memory that internally stores instructions for configuring the processor to acquire a bitstream, An audio decoder device comprising, wherein the bitstream is The encoded version of the audio signal, The first dynamic range control, i.e., DRC, gain sequence, is determined by encoder-side processing that applies the aforementioned audio signal to a first DRC characteristic. An index of the first DRC characteristic, wherein the index identifies or indicates the first DRC characteristic, Multiple instances of the integrated loudness update value over time, Includes, The bitstream includes encoder-side DRC normalization gain, and the processor cancels out the encoder-side DRC normalization gain when applying decoder-side DRC loudness normalization. Audio decoder device.

18. Processor and A memory that internally stores instructions for the processor to generate a bitstream, An audio decoder device comprising, wherein the bitstream is The encoded version of the audio signal, The first dynamic range control, i.e., DRC, gain sequence, is determined by encoder-side processing that applies the aforementioned audio signal to a first DRC characteristic. The index of the first DRC characteristic, Multiple instances of the integrated loudness update value over time, Includes, The bitstream controls the decoder-side processing, specifically how it applies DRC to the audio signal while performing loudness normalization. Audio decoder device.

19. The audio encoder device according to claim 18, wherein the processor inserts a flag into the bitstream, and when the flag has a first value, the bitstream includes encoder-side loudness normalization gain, or when the flag has a second value, the bitstream does not include encoder-side loudness normalization gain.

20. The audio encoder device according to claim 19, wherein when the flag has the first value, the loudness normalization gain is applied by the encoder-side processing when determining the first DRC gain sequence.

21. The acquisition of a bitstream, wherein the bitstream includes an encoded version of an audio signal, a first dynamic range control (DRC), a gain sequence, an index of the first DRC characteristic, wherein the index identifies or indicates the first DRC characteristic, and a plurality of instances over time of an integrated loudness update value. The inverse DRC characteristics are obtained using the aforementioned index, The process involves applying the inverse DRC characteristics to the first DRC gain sequence, followed by loudness normalization to generate a normalized loudness sequence. The normalized loudness sequence is applied to a second DRC characteristic to generate a second DRC gain sequence, The second DRC gain sequence is applied to the audio signal to generate compressed audio, Digital audio methods, including [specific method / technique].

22. The bitstream includes loudness normalization gain applied by the encoder when determining the first DRC gain sequence by applying the audio signal to the first DRC characteristics, The bitstream instructs the processor to perform loudness normalization by using the loudness normalization gain in the bitstream to cancel out or reverse the loudness normalization applied by the encoder when determining the first DRC gain sequence. The method according to claim 21.

23. The method according to claim 21, wherein when the bitstream includes a flag and the flag has a first value, the first DRC gain sequence is determined by the encoder-side processing which applies the audio signal to the first DRC characteristics with loudness normalization.

24. The method according to claim 23, wherein when the flag has a second value, the first DRC gain sequence is determined by the encoder-side processing that applies the audio signal to the first DRC characteristics without loudness normalization.

25. Performing loudness normalization is Adjust the normalized loudness sequence, and then apply the adjusted loudness sequence to the second DRC characteristic. The method according to any one of claims 21 to 24, including the method described in any one of claims 21 to 24.

26. Encoding the audio signal to generate an encoded version of the audio signal, The process involves processing the aforementioned audio signal to generate multiple instances of the integrated loudness update value over time, The aforementioned audio signal is applied to the dynamic range control, i.e., DRC, characteristics, to determine the DRC gain sequence. The process includes generating a bitstream containing the encoded version of the audio signal, the DRC gain sequence, the index of the DRC characteristic, and the multiple instances of the integrated loudness update value over time, wherein the bitstream controls the manner in which the decoder-side processing applies DRC to the audio signal while performing loudness normalization. Digital audio processing.

27. The further comprising inserting a flag into the bitstream, wherein the bitstream includes encoder-side loudness normalization gain when the flag has a first value, or does not include encoder-side loudness normalization gain when the flag has a second value. The process described in claim 26.

28. Processor and A memory that stores and holds instructions internally, An audio decoder device comprising the following, wherein the instruction causes the processor to A bitstream is obtained, the bitstream including an encoded version of the audio signal, a momentary loudness sequence, and an integrated loudness value. The normalization gain is calculated by combining the DRC input loudness target with the integrated loudness value extracted from the bitstream. The instantaneous loudness sequence extracted from the bitstream is adjusted using the loudness normalization gain to generate a normalized instantaneous loudness sequence. The DRC gain sequence is generated by applying the normalized instantaneous loudness sequence to the DRC characteristics. By applying the DRC gain sequence to the audio signal, DRC is performed on the audio signal. An audio decoder device configured in such a way.

29. The audio decoder device according to claim 28, wherein the instantaneous loudness sequence in the bitstream is not loudness normalized.

30. The audio decoder device according to any one of claims 28 to 29, wherein the integrated loudness value is one instance of a plurality of instances of integrated loudness update values ​​included in the bitstream, and adjacent instances are separated by, for example, 1 to 10 seconds, and the integrated loudness update value represents the moving average integrated loudness of the audio signal.

31. The audio decoder device according to any one of claims 28 to 29, wherein the bitstream is a file on which the integrated loudness value is written along with the instantaneous loudness sequence and the encoded version of the audio signal.

32. Processor and A memory that internally stores instructions for the processor to generate a bitstream, An audio decoder device comprising: the bitstream includes an encoded version of an audio signal; an instantaneous loudness sequence of the audio signal; and instructions for controlling the manner in which the decoder applies dynamic range control, i.e., DRC, to the audio signal with loudness normalization when applying the instantaneous loudness sequence to the DRC characteristics. Audio encoder device.

33. The process involves obtaining a bitstream, wherein the bitstream includes an encoded version of the audio signal, a momentary loudness sequence, and an integrated loudness value. The normalization gain is calculated by combining the DRC input loudness target with the integrated loudness value extracted from the bitstream, The instantaneous loudness sequence from the bitstream is adjusted using the loudness normalization gain to generate a normalized instantaneous loudness sequence. The DRC gain sequence is generated by applying the normalized instantaneous loudness sequence to the DRC characteristics. Applying the DRC gain sequence to the audio signal performs DRC on the audio signal. Digital audio processing, including.

34. Encoding the audio signal to generate an encoded version of the audio signal, The process of the aforementioned audio signal to generate an instantaneous loudness sequence of the aforementioned audio signal, A bitstream is generated that includes the encoded version of the audio signal, the instantaneous loudness sequence, and instructions for the decoder to control the manner in which dynamic range control, i.e., DRC, is applied to the audio signal with loudness normalization when the instantaneous loudness sequence is applied to the DRC characteristics. Digital audio processing, including.