Retarded loudness adjustment for dynamic range control

By delaying loudness normalization processing at the decoding end, the inaccuracy of dynamic range control in live audio is solved, achieving more accurate dynamic range compression and improving audio quality and experience.

CN114464199BActive Publication Date: 2026-04-14APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2021-11-09
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In live audio streaming and recording, existing technologies struggle to accurately adjust dynamic range compressor characteristics because the program loudness is unknown before encoding, leading to inaccurate loudness normalization and pumping artifacts.

Method used

The loudness normalization process is postponed from the encoding end to the decoding end. By estimating and applying the dynamic range control gain sequence at the decoding end, and using smoothing filters and delay blocks to process the audio signal, combined with the MPEG-D DRC standard extension, loudness normalization at the encoding end is supported.

Benefits of technology

It improves the accuracy of dynamic range control, reduces loudness shift, and enhances the audio playback experience, making it suitable for live streaming and recording to files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114464199B_ABST
    Figure CN114464199B_ABST
Patent Text Reader

Abstract

A bitstream is obtained by a decoding end, which includes an encoded version of an audio signal and a sequence of instantaneous loudnesses of the audio signal. The sequence of instantaneous loudnesses has not been loudness normalized. A dynamic range control (DRC) gain sequence is generated by applying the sequence of instantaneous loudnesses to a DRC characteristic with loudness normalization. The DRC gain sequence is applied to the decoded audio signal. Other aspects are also described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Audio programs, such as music, podcasts, short video clips from live recordings, or feature films, have loud and soft segments that limit their dynamic range and dynamic range. In many cases, such as listening through headphones in a noisy environment or through home speakers in a nighttime scene, it is desirable to reduce the dynamic range and dynamic range of the played sound to improve the listener's experience. Dynamic range compressors are used for this purpose. They are digital signal processors that apply time-varying gain to the input digital audio signal (of the audio program) to amplify soft segments and attenuate loud segments of the audio signal. To avoid audible pumping artifacts that can be caused by dynamic range compression of the audio signal, a loudness normalization process is performed that "aligns" the input audio signal to a compression feature or profile while compressing the audio signal according to that feature. This can be done by canceling the instantaneous loudness of the input audio signal with its program loudness, where program loudness is a calculated value designed to describe the overall loudness of the audio program (also known as integrated loudness). Summary of the Invention

[0002] Audio coding standards define methods for dynamic range compression that generate Dynamic Range Control (DRC) gains at the encoding end, where an audio program is being created, prepared for distribution, or stored / archived. This DRC gain, referred to herein as a DRC gain sequence, is time-aligned with its associated audio program such that one or more gain values ​​in the sequence will be applied to the corresponding digital audio frames of the audio program. This DRC gain sequence is then formatted, for example, as metadata associated with the audio program, into one or more bitstreams. If needed at the decoding end (typically during playback of the decoded audio signal), the decoding end acquires the bitstream and applies the in-stream DRC gain to compress the dynamic range of the decoded audio signal. The advantage of such metadata-based methods is improved quality because the lead-in time interval for offline encoding of the DRC gain can be larger compared to real-time compression. Another advantage is that, for example, compression characteristics can be controlled at the encoding end based on the expertise of the audio program creator or distributor.

[0003] For metadata-based DRC in online applications (e.g., live audio streaming and recording live audio to files), there are challenges if the program loudness of the audio program being streamed or written to a file for storage is not yet known (because the audio program has not ended). This is because if the actual program loudness of the audio program (which can only be determined after the audio program has ended) deviates significantly from the expected or predicted program loudness, the compressor characteristics may not be properly adjusted (or loudness normalized).

[0004] Some aspects of this disclosure describe novel digital signal processing methods that postpone loudness adjustment (loudness normalization) for dynamic range control (DRC) from the encoder to the decoder. Other aspects describe techniques for altering compressor characteristics at the decoder using loudness normalization when performing metadata-based DRC gain sequence processing. These aspects are particularly advantageous for applications such as live streaming and live recording to files.

[0005] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed descriptions below and specifically pointed out in the claims section. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description

[0006] The aspects of this disclosure are illustrated by way of example and are not limited to the illustrations in the accompanying drawings, in which similar reference numerals indicate similar elements. It should be noted that references to “a” or “an” aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a given drawing may be used to illustrate more than one aspect of this disclosure, and for a given aspect, not all elements in that drawing may be necessary.

[0007] Figure 1 An example of a DRC characteristic curve is shown.

[0008] Figure 2 This is a block diagram of an audio codec system that applies DRC at the decoding end and does not perform loudness normalization at the encoding end.

[0009] Figure 3 This is a block diagram of an audio codec system suitable for live streaming that applies DRC at the decoding end and does not perform loudness normalization at the encoding end.

[0010] Figure 4 This is a block diagram of an audio codec system that applies DRC at the decoding end and does not perform loudness normalization at the encoding end, suitable for live recording for storage or archiving.

[0011] Figure 5 This describes a part of an MPEG-D DRC compliant audio codec system that applies DRC at the decoding end.

[0012] Figure 6 This describes a part of an MPEG-D DRC compliant audio codec system that applies DRC at the decoding end and loudness normalization at the encoding end.

[0013] Figure 7This describes a part of an MPEG-D DRC-compliant audio codec system that applies DRC and loudness normalization at the decoding end.

[0014] Figure 8 It is a flowchart of the process for generating new encoding end processes that can generate backward-compatible and non-backward-compatible MPEG-D DRC bitstream extensions.

[0015] Figure 9 This is a flowchart of a new decoding process that can use backward-compatible or non-backward-compatible MPEG-D DRC bitstream extensions to generate DRC gain sequences.

[0016] Figure 10A and Figure 10B It is a block diagram of an audio codec system conforming to MPEG-D DRC, in which a backward-compatible encoder produces a backward-compatible bitstream that is processed by both the new decoder and the traditional decoder. Detailed Implementation

[0017] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Where the shape, relative position, and other aspects of the described components are not explicitly defined, the scope of the invention is not limited to the components shown, which are for illustrative purposes only. Furthermore, while many details have been set forth, it should be understood that some aspects of this disclosure can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0018] To properly apply dynamic range control to an audio signal, the compressor characteristics (DRC characteristics, DRC profile) should be "aligned" with the loudness level range of the audio signal. For example, refer to... Figure 1 Aligned along the input horizontal axis, the zero-crossing point of the DRC characteristic curve is approximately located at the center of the loudness level range of the audio signal. The level at the zero-crossing point is also known as the DRC input loudness target—in Figure 1In the example set of characteristic curves shown, this level is approximately -31 dB. The center of this loudness level range could be, for example, the average level of a sound program, or the average dialogue level within a sound program. The process of achieving this alignment is referred to here as loudness normalization, which combines the DRC of the audio signal to achieve a given loudness target. For example, the loudness of an audio signal (sound program) could be a single value called integrated loudness. Integrated loudness is a measure of the loudness of an audio signal that is similar to the root mean square (RMS) but more realistic in terms of human hearing. Integrated loudness can be equated to program loudness because it measures the loudness of a sound program over its entire duration. To achieve loudness normalization, when the integrated loudness is given in decibels (dB), the integrated loudness can be subtracted from the DRC input loudness target to derive a normalized gain in dB. This normalized gain is added to the output of a loudness model that calculates the instantaneous loudness of the audio signal (sound program). The instantaneous loudness can be a sequence of loudness values, each calculated based on a corresponding digital audio frame that constitutes the input digital audio signal (and represents human-perceived loudness). Another way to achieve loudness normalization is to... Figure 1 The DRC characteristic curve depicted in the figure shifts to the right or left of zero dB (determined by the magnitude of the normalized gain). Figure 1 In the example, the curve has been shifted to the left to -31dB (the loudness target in this example), and is therefore correctly aligned with a sound program with an integrated loudness of -31dBA (A-weighted) or LKFS (loudness K-weighted level full scale) (and can therefore be directly applied to that sound program) - in other words, the normalized gain in this case will be zerodBA.

[0019] If the integrated loudness of the audio program is still unknown while the dynamic range control signal processing is in progress, as is the case with live audio, then prediction is needed to apply loudness normalization. However, this prediction may be incorrect, resulting in an undesirable bias in the DRC gain, or a pumping effect in the DRC gain—an undesirable loudness shift between the uncompressed and compressed portions of the audio signal.

[0020] To reduce the possibility of unwanted loudness shift, one aspect of this disclosure applies loudness normalization to the decoding end of an audio codec system or method rather than to DRC at the encoding end. Figure 2 The hardware block diagram shows an example of an audio codec system and associated methods. Various hardware blocks of this audio codec system and methods can be implemented by a programmable processor. In this method, integrated loudness (required in conjunction with loudness normalization performed by DRC for playing or archiving / storing the decoded audio signal) can be obtained in at least two instances, which will be discussed below. Figure 3 and Figure 4 Describe it.

[0021] from Figure 2 Initially, the audio codec system has an encoding end, which can be implemented by one or more processors that execute or are configured by instructions stored in memory, often referred to herein as a "programming processor," such as in one or more servers. The upper audio signal processing path includes sidechains that calculate or estimate the instantaneous loudness of the digital input audio signal (sequence) (also referred to herein as a sound program). This estimation is based on a perceptual loudness scale (e.g., the acoustic scale), and is therefore approximately logarithmic. To smooth this instantaneous loudness sequence over time, a smoothing filter, as shown in the figure, can be applied. It produces smoothness in regions of the input audio sequence where compression gain changes are not required but macroscopic dynamic loudness transitions remain unaffected.

[0022] The lower audio signal processing path includes a delay block that delays the input audio sequence to compensate for the latency caused by the sidechain. The smoothed loudness sequence and the delayed input audio sequence are then fed to the encoder.

[0023] The encoder performs a bitrate reduction operation on one or both of its inputs and can produce one or more bitstreams containing bitrate-reduced versions of one or both of its inputs. These bitstreams can then be transmitted to the decoder (e.g., via the internet), or they can be written to a file for storage or archiving until accessed by the decoder process. Smooth loudness (referring to a sequence of smooth loudness values, or a single smooth loudness value) can be carried as metadata in the same bitstream as the delayed input audio sequence, for example, associated with its “corresponding” high-level audio codec, AAC, or audio frame. This is also referred to as being at the audio layer. Alternatively, the smooth loudness sequence and other loudness values, such as integrated loudness updates and DRC payloads (discussed further below), may not be transmitted at this audio layer, but at a higher layer, such as at the file format level. In both cases, one or more bitstreams are produced, providing encoded audio along with associated metadata, such as the smooth loudness sequence, or, as described elsewhere below, instructions for applying the encoder source DRC gain sequence to the decoder.

[0024] The decoding end can also be implemented as a programmable processor, such as one or more processors that execute instructions stored in memory or configured as part of an audio playback device by instructions stored in memory. It should be noted that the decoding process can be implemented in the same audio playback device as the encoding process. Alternatively, the decoding process can be implemented in an audio playback device that is separate from the program processor executing the encoding process. Examples of such audio playback devices include smartphones, tablets, digital media players, headphones, or vehicle infotainment systems. At the decoding end, the decoder undoes the encoder's bitrate reduction operation to recover the smoothed loudness sequence and the delayed input audio sequence. The decoded smoothed loudness values ​​are then mapped to "corresponding" DRC or compression gain values. This mapping is implemented, for example... Figure 1 The diagram shows a storageless input-output function of one of the curves (or alternatively, any other desired curve). This mapping constitutes a compressor characteristic or compression profile (DRC characteristic) whose output is a time-varying gain (a sequence of DRC gain values) that is a function of the time-varying input loudness level. The mapping may also include a conversion from the logarithmic loudness domain to the linear domain of the DRC gain. If compression is required, the DRC gain values ​​(sequence) are applied to the decoded audio signal, as indicated by the multiplier symbol in the diagram. Although not shown, the compressed audio can be passed to a playback processing block, which ultimately produces transducer (speaker) driver input signals that convert the compressed audio into sound.

[0025] from Figure 2 As can be seen, the smoothed loudness sequence is adjusted or normalized at the decoding end before being input to the DRC mapping block. For example, a constant integrated loudness (a single value) (in the dB domain) can be subtracted from the DRC input loudness target to derive a normalized gain in dB. This normalized gain is added to each smoothed loudness value in the smoothed loudness sequence to produce the normalized loudness sequence used in the DRC process. This DRC process has at least two applications, such as live or real-time streaming and live recording to files for storage or archiving.

[0026] In one such application, now refer to Figure 3The input audio to the encoder is a live or real-time digital audio recording that is being transmitted to the decoder, for example, via the internet. This input audio contains audio captures of live or real-time events occurring simultaneously with encoding and bitstream transmission. Thus, a single integrated loudness value representing the entire sound program cannot be calculated until the live event ends. Before this, the integrated loudness measurement block in the encoder collects samples of the live audio over a time interval longer than a single audio frame (e.g., several seconds), which is then sent to a time-aligned encoder, and a loudness metric for that interval is calculated. It then "integrates" or collects several such metrics back to the beginning of the sound program, for example, by averaging them, to calculate an integrated loudness update. This integrated loudness update can be a measure of integrated loudness only for the portion of the sound program that has been played or streamed before the current update. For example, this measurement can be repeated periodically to effectively produce a "running average" integrated loudness, and the latest integrated loudness update (which is a single value) is transmitted to the decoder. It should be noted that the term "running average" used here does not require performing an actual averaging; it only requires performing some measure of the loudness of the sound program from the start of the program to the current update, based on collected loudness measurements (including statistical results evaluating the collected loudness measurements). An update (running average) can be calculated and then provided as part of a bitstream that also contains the encoded sound program (encoded audio signal), as multiple instances of an integrated loudness update field, where adjacent instances in the bitstream are spaced one to ten seconds apart over the duration of the sound program.

[0027] It should also be noted that the term "integrated loudness update" can also be referred to as running average loudness or "partial integrated loudness"; at the end of the audio program, the final or final integrated loudness update can represent the loudness of the entire audio program (also known as integrated loudness or program loudness, for example, the example of measuring audio program loudness and true peak audio level described in the algorithm of ITU-R BS.1770-4 (10 / 2015) recommendation).

[0028] At the decoding end, the decoder acquires the bitstream and extracts an integrated loudness update from it. The decoding process then applies this integrated loudness update to the loudness normalization of the DRC process. This can be accomplished, for example, by adding a single loudness normalization gain value (e.g., the difference between the DRC input loudness target and the integrated loudness update value) to the decoded or recovered instantaneous loudness sequence before inputting it to the DRC feature mapping block. Alternatively, loudness normalization can be accomplished by shifting the DRC feature along its input axis by the amount of a loudness normalization gain value. This loudness normalization gain can be updated periodically during the transmission of the bitstream (audio program), where the most recent partial integrated loudness value (integrated loudness update) has already been calculated at the encoding end for the passed portion of the live event.

[0029] In another application, now refer to Figure 4 The input audio to the encoder is a live or real-time digital audio recording of an event, which is written to a file for archiving or storage at the end of the recording (when the event ends). A single integrated loudness value representing the program loudness of the entire live audio event is calculated by the integrated loudness model block at the end of the recording and provided to the encoder as soon as the event ends. The encoder writes this integrated loudness value, along with an encoded version of the live audio and an encoded version of the instantaneous (and smoothed) loudness sequence calculated by the loudness model (based on the same live audio), to the file. At the decoder, the decoder acquires the file (bitstream) and decodes the input audio and the instantaneous loudness sequence from the file, extracting the integrated loudness value from the file. The decoding process then uses this integrated loudness value to loudness normalize the decoded instantaneous loudness sequence before inputting it to the DRC (compression) mapping block, and then applies the output of the DRC (compression) mapping block to the decoded input audio during playback (if compression is required).

[0030] On the one hand, this smoothing filter is a nonlinear filter, such as the nonlinear filter described in U.S. Patent No. 10,109,288. A useful property of this filter is that its output can undergo the same amount of level shift as the input. This means that when f(x) is defined as a nonlinear filter function, x(n) is defined as the input signal, and y(n) is defined as the output, it is possible to write...

[0031] y(n)=f(x(n))

[0032] Given an input signal shifted by ΔL, if the output is shifted by ΔL, then f(x) satisfies the shift property, or can be expressed mathematically as:

[0033] y(n)+ΔL=f(x(n)+ΔL)

[0034] This is beneficial because it avoids any sidechain processing in the encoding end that depends on the absolute loudness value.

[0035] Another aspect of this disclosure describes the application of DRC in accordance with the MPEG-D DRC standard ISO / IEC "Information Technology - MPEG Audio Technology - Part 4: Dynamic Range Control", ISO / IEC 23003-4:2020 ("MPEG-D DRC"), extending it to support loudness normalization at the encoder end. Figure 5 A simplified block diagram is shown for a portion of the MPEG-D DRC process that generates and applies the DRC gain based on the metadata decoding gain in the bitstream obtained from the encoding end. MPEG-D DRC provides predefined DRC features and a flexible way to encode parameterized features.

[0036] exist Figure 5 In the encoding process, the smoothed instantaneous loudness sequence (calculated for the input audio sequence) is applied to the selected DRC feature (also known as a "mapping block," as described above). Figure 2 (The DRC feature map block's output produces a DRC gain sequence, which is then fed to the DRC encoder. The encoder performs bitrate downsampling to encode its input sequence into one or more bitstreams, which are then transmitted to or otherwise provided to the decoder. At the decoder, the DRC decoder undoes the bitrate downsampling to recover the DRC gain sequence (the decoded DRC gain sequence). This decoded DRC gain sequence is then applied to the decoded audio signal (if compression is required).

[0037] MPEG-D DRC also supports a type of decoding-end processing that removes the DRC features used to compress audio programs from... Figure 5 The DRC features used by the encoder (to calculate the DRC gain sequence inserted into the bitstream as metadata) are changed to different DRC features that can be processed by the decoder based on the current playback or listening conditions. This is as follows: Figure 6 The implementation in the decoder shown first applies the encoder's DRC gain sequence to inverse feature A, which is the inverse of the DRC feature A applied at the encoder to produce the encoder's DRC gain sequence. An index (identifier or pointer) of this DRC feature A (used at the encoder to produce the DRC gain sequence) can be provided in the bitstream so that the decoder can recognize this inverse feature A. Applying the DRC gain sequence as input to inverse feature A produces a recovered, smooth, instantaneous loudness sequence. If quantization effects are ignored, this recovered loudness sequence (at the output of the inverse feature A block) is essentially the smooth loudness sequence used by the encoder. Therefore, the recovered loudness sequence can be applied to a second DRC feature B to produce a second DRC gain sequence (more suitable than DRC feature A) that is better suited for compressing the decoded audio signal. This second DRC gain sequence is then applied to the decoded audio (if, for example, compression is required during playback).

[0038] Based on one aspect of the content disclosed in this article, using Figure 2 The method shown replaces Figure 6 The loudness normalization of the sidechain in the encoder is shown. This means that the integrated loudness-based offset (normalized gain) is applied to the decoder, not the encoder. Figure 7A block diagram of this system is shown. This is also referred to herein as an enhanced MPEG-D DRC-compatible system (hereinafter also referred to as having a “new” encoder and a “new” decoder). Such a system has a block at its encoding end called an integrated loudness measurement, the output of which provides an update of the integrated loudness value, as combined above. Figure 3 The discussion focuses on providing integrated loudness updates to an audio encoder. Here, the encoder is a DRC encoder, which also encodes a DRC gain sequence (in addition to the input audio). This DRC gain sequence can be combined... Figure 6 The determination is performed as described above. The encoded DRC gain sequence and the integrated loudness update are provided to the decoder via one or more bitstreams. The DRC gain sequence can be formatted as metadata and associated with the encoded input audio, which is also provided to the decoder.

[0039] The integrated loudness measurement is a running measurement of integrated loudness (also referred to herein as the running average), which begins at the start of the audio program and continues to "integrate" the audio signal of the audio program over time, with the aim of calculating only the integrated loudness value of the portion of the audio program that has passed. As the audio signal (audio program) continues, the integrated loudness measurement is updated, for example periodically (e.g., every ten seconds). These integrated loudness updates are written to the bitstream (e.g., by the DRC encoder). This can be supported in MPEG-D DRC by writing the updates to an extended field or extended payload of the audio bitstream, or by writing them as part of a separate metadata track in an MP4 file. Without introducing additional system latency, the predictability of this update is equivalent to the latency of the sidechain that produces the DRC gain sequence (at the output of DRC feature block A). Greater predictability improves the first integrated loudness update at the start of the audio program, meaning it can be closer to the program loudness of the audio program.

[0040] in available Figure 7 In the first scenario shown, the input audio is live audio, which is simultaneously streamed to the decoder via a bitstream (e.g., over the internet). In this case, program loudness cannot be provided during streaming (because the live audio event has not yet ended). In this case, the DRC (applied to the decoder) is dynamically adjusted based on the integrated loudness update in the stream, or loudness normalization is performed—a dynamically changing normalized gain that can be equal to the difference between the DRC input loudness target value and the dynamically changing integrated loudness update. To limit the rate of change of the integrated loudness update, the update sequence can be smoothed at the beginning of the stream rather than at the end. Additionally, the initial update value (at the beginning of the stream) can take into account the expected loudness of the input audio. For example, this expected loudness might be the result of careful setup in a professional studio and experimental measurements of the initial portion of the input audio.

[0041] In the second case, the input audio (at the encoding end) is as follows: Figure 4 The image shown is a live audio recording being written to an audio file at the encoding end (not a live stream). In this case, the final integrated loudness update (the true integrated loudness or program loudness of the sound program) can be written to the file at the end of the recording without rewriting the file. When attempting to comply with MPEG-D DRC, this is achieved by writing the final integrated loudness update (at the encoding end) to a loudness "box" or field at the ISO base media file format level. This audio stream loudness box type is called a ludt. See still for reference. Figure 7 When the decoder obtains the encoded audio and its associated encoder-side DRC gain sequence and integrated loudness update, the decoder process can apply DRC by determining the DRC gain sequence based on the loudness-normalized version of the decoded audio signal (using DRC feature B). In this example, this normalization adjusts the smooth instantaneous loudness recovered at the output of inverse feature A, preferably achieved by using the final integrated loudness update value written to the loudness box. If the recording ends at the encoder without adding the loudness box to the stream, DRC with loudness normalization can still be applied at the decoder using the loudness update integrated in the stream.

[0042] Because the integrated loudness update in the stream can change slowly over time (e.g., every 1 to 10 seconds), normalization will effectively cause the DRC feature B to shift accordingly. When the integrated loudness update is based on a short duration (the elapsed time interval) of the audio program, this shift may become audible during playback of the decoded audio at the start of recording or streaming. To limit the rate of change of the integrated loudness update, the update itself may be smoothed at the start of recording or streaming rather than at the end.

[0043] According to Figure 6 During the encoding process, where the input audio is a live recording to a file, it can be compressed using sidechain loudness normalization (DRC), then encoded at the encoding end and written to the file. This may inherently result in compressed audio output, or, if not inherently the same as the output during the decoding process, it can be compared to... Figure 7 Compared to the output shown, in Figure 7 In this process, the decoded audio is compressed at the decoding end, and loudness normalization is based on an integrated loudness update that includes the bitstream. However, as... Figure 7 As shown, the advantage of deferring loudness normalization to the decoding end is that it improves the listening experience when playing files by adding a final integrated loudness update at the ISO base media file format level (MP4 level) only after recording or the event has ended.

[0044] Now go to Figure 8This is a flowchart of the new encoding-side process, which generates backward-compatible and non-backward-compatible MPEG-D DRC bitstream extensions for the decoding-side DRC. The backward-compatible bitstream extension fields or payload can be processed by a traditional decoder (decoding-side process) to perform DRC based on the extension, but loudness normalization is not performed (when DRC is applied to the decoded audio signal). For an example of such a traditional decoder, see [link to example]. Figure 6 Non-backward-compatible bitstream extensions cannot be processed by this traditional decoder (to produce compressed audio). This dual feature can be enabled as follows.

[0045] A flag included in this bitstream can be defined, for example, `characteristicV1Override`. The encoder can set or clear this flag, as described below. To produce a backward-compatible bitstream, the flag is assigned a first value, for example, `characteristicV1Override = 1`, and in this case, the bitstream will also include a loudness normalization gain, also known as `encDrcNormGainDb`. In this mode, the encoder process determines the first DRC gain sequence by applying the loudness normalization gain (here also referred to as the encoder-side DRC normalization gain) to the audio signal using a first DRC characteristic with loudness normalization. See also Figure 10A and Figure 10B This is a block diagram of an MPEG-D DRC compliant audio codec system, where a backward-compatible encoder produces a backward-compatible bitstream processed by both the new decoder and the legacy decoder. In the case of a live recording of the input audio, an integrated loudness update is also calculated and provided to the encoder (to be incorporated into the bitstream). Figure 10A As shown, the loudness normalization gain can be calculated by subtracting the predicted program loudness value (e.g., assuming it is in dBA) from the DRC input loudness target.

[0046] The loudness normalization gain `encDrcNormGainDb` is a value applied during the new backward-compatible coding process to produce a backward-compatible bitstream, where the DRC gain sequence is obtained using loudness normalization (to DRC feature A). This bitstream can be processed by both new and legacy decoders, for example... Figure 10BAs shown. When this bitstream is processed by a traditional decoder, the latter does not apply loudness normalization during DRC. When this bitstream is processed by a new decoder that applies loudness normalization during DRC, encDrcNormGainDb is used to compensate for, neutralize, or undo the application of encDrcNormGainDb by the backward-compatible encoder, so that a more accurate loudness normalization can be applied using the integrated loudness update. In other words, when applying decoder-side DRC loudness normalization, the processor of this new decoder compensates for the encoder-side DRC normalization gain.

[0047] Back Figure 8 To enable the bitstream to be processed by both traditional and new decoders, when the flag has the first value, such as characteristicV1Override = 1, the backward-compatible bitstream may also include a first DRC configuration field, such as UNIDRCCONFEXT_V1, and a second DRC configuration field, such as UNIDRCCONFEXT_V2. The first DRC configuration field instructs the decoding process to apply DRC to the decoded audio signal without loudness normalization, for example, as... Figure 10B The traditional decoder block is shown in the diagram. The second DRC configuration field instructs the decoding process to apply DRC to the decoded audio signal with loudness normalization, for example, as... Figure 10B The new decoder block is shown in the diagram.

[0048] Still referencing Figure 8 This new encoder can create a non-backward-compatible MPEG-D DRC bitstream extension (a bitstream extension that cannot be processed by a traditional decoder to produce compressed audio), as described below. Note that the encoder may wish to do so if it is aware that only the new decoder will process the bitstream. In such a bitstream, the flag has a second value, for example, characteristicV1Override = 0, and the bitstream does not include loudness normalization gain (which is intended for use by the decoder). Furthermore, the first DRC configuration field, such as UNIDRCCONFEXT_V1, is also omitted from the bitstream. Figure 9 A new decoder capable of handling such bitstreams is shown. In other words, when the flag has the second value, for example, characteristicV1Override = 0, the bitstream includes the second DRC configuration field but not the first DRC configuration field.

[0049] Figure 9This is a flowchart of a new decoder process that can use backward-compatible or non-backward-compatible MPEG-D DRC bitstream extensions to generate a new DRC gain sequence. The process may begin by parsing the bitstream to detect the second DRC configuration field, such as UNIDRCCONFEXT_V2, and the flag characteristicV1Override. In response to the flag having a first value (e.g., characteristicV1Override = 1), the process applies DRC to the audio signal, for example, as... Figure 10B As shown (the new decoder block), using the DRC feature B and loudness normalization, including using i) the loudness normalization gain (e.g., encDrcNormGainDb) and ii) multiple instances of integrated loudness updates (both of which are obtained by the DRC decoder along with the audio signal from the acquired bitstream).

[0050] On the one hand, still refer to Figure 9 When the flag has the first value, such as characteristicV1Override = 1, the index of the first DRC feature that can be included in the first DRC configuration field is overwritten by the index of the first DRC feature included in the second DRC configuration field. For example, the MPEG-D DRC can define DRC features 1 to 6 (also referred to herein as traditional index values ​​or traditional ranges) that can be recognized by a conventional MPEG-D DRC decoder. In this disclosure, according to the enhanced MPEG-DDRC process, those same features are copied with different index values ​​(also referred to herein as new index values ​​or new ranges), such as 65 to 70. In other words, the conventional feature can be referenced by its conventional indices 1 to 6 or by its new indices 65 to 70; the parameters of the conventional features are the same, as shown in the table below.

[0051] Table 6 - Parameters of DRC characteristics with index ranges of 1 to 6 and 65 to 70

[0052]

[0053] When the new encoding process generates a backward-compatible bitstream (characteristicV1Override=1), Figure 8When the flowchart on the right is executed, both the first (V1) DRC configuration extension field and the second (V2) DRC configuration extension field are generated. The first DRC configuration field refers to one or more of the traditional indices 1 to 6, but not any of the new indices 65 to 70, to achieve backward compatibility with traditional decoders. The V2 extension field may refer to one or more of the new index values, or it may refer to one or more of the traditional index values. The new index value effectively signals to the new decoder (a decoder conforming to the enhanced MPEG-D DRC process of this disclosure) that loudness normalization may be required when generating the second DRC gain sequence. Only the UNIDRCCONFEXT_V2 extension supports DRC feature indices 65 to 70, which require loudness normalization in the decoder.

[0054] The new decoding process can decode both the V1 and V2 extended fields, such as... Figure 9 As shown on the right, this extracts two indices (two different index values) pointing to the same DRC feature. In this case, the V2 index is said to overwrite the V1 index because characteristicV1Override=1, and the new decoder will replace the DRC feature index obtained from the UNIDRCCONFEXT_V1 extension with the DRC feature index from the UNIDRCCONFEXT_V2 extension.

[0055] Back Figure 8 When generating a non-backward-compatible bitstream (to provide to a new decoder, not a legacy decoder), the flag `characteristicV1Override` is set to zero, and the `UNIDRCCONFEXT_V2` extension is generated in that bitstream. This `UNIDRCCONFEXT_V2` extension contains essentially the same bitstream fields as the `UNIDRCCONFEXT_V1` extension. While `UNIDRCCONFEXT_V1` does not support characteristics 65 to 70, the transmitted `UNIDRCCONFEXT_V2` does. Since loudness normalization of the DRC sequence generated at the encoder is not applied in this case—see... Figure 7 - It was not compensated in the decoder (see also) Figure 7 This situation is equivalent to... Figure 10BDuring the decoding process, the normalization gain, for example, encDrcNormGainDb, is set to zero. When this bitstream is parsed by this new decoding process, in response to i) a flag having a second value and ii) an index of a first value (e.g., in the range of 65 to 70), the decoding process uses a second DRC feature B and applies DRC to the audio signal using loudness normalization with integrated loudness updates but without using loudness normalization gain (e.g., the value of encDrcNormGainDb in the summation block is set to zero). In other words, when the normalized loudness sequence is generated at the input of DRC feature B, encDrcNormGainDb is set to zero.

[0056] However, if the new decoder encounters i) a flag with a second value and ii) an index that is a second value different from the first value (e.g., in the range 1 to 6), the decoding process applies DRC to the audio signal (using the second DRC feature B) but does not perform loudness normalization. In other words, referencing Figure 10B At the output of the inverse characteristic A, the recovered smooth instantaneous loudness sequence is not adjusted (before being input to the DRC characteristic B) - therefore, the summation block shown in the figure does not exist.

[0057] The following appendix includes a draft specification of the method for delayed loudness normalization proposed within the framework of the MPEG-D DRC standard. This document includes an efficient method for generating bitstreams with new information that can also be decoded using conventional decoders.

[0058] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that these aspects are merely illustrative and not limiting of the invention, and that the invention is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.

[0059] appendix:

[0060] International Organization for Standardization

[0061] International Organization for Standardization

[0062] ISO / IEC JTC1 / SC29 / WG6, MPEG audio encoding

[0063] ISO / IEC JTC1 / SC29 / WG6 Mxxxx

[0064] Month of 2020

[0065] The title proposes revisions to ISO / IEC 23003-4 to improve support for real-time coding.

[0066] SOURCE Frank Baumgarte, Apple Inc.

[0067] 1 Introduction

[0068] MPEG-D DRC (ISO / IEC 23003-4) is a codec-independent tool for loudness and dynamic range control based on metadata. When the dynamic range compressor in the encoder generates DRC gain metadata, the audio signal is typically loudness normalized at the compressor's input. This normalization is used to align the DRC characteristics with the loudness profile of the source signal. In real-time encoding scenarios, such as live broadcasts, the loudness profile of the source signal may be unknown at the start of encoding, and DRC alignment may be disabled.

[0069] 2. Problem Description

[0070] The MPEG-D DRC standard recommends a two-pass encoding method. In the first pass, the overall loudness of the audio content is measured. In the second pass, the DRC input signal is normalized based on the loudness measurement, and metadata is encoded. This method is not suitable for real-time encoding or real-time applications. Therefore, loudness measurement cannot be performed before the encoding pass begins. One workaround might be to calibrate the audio frequently used in professional studios, which provides a fairly predictable production loudness that can serve as an estimate of the measurement. In other environments, the loudness of the audio signal is less predictable and can therefore deviate significantly from the intended value used to set the encoder. Such deviations can lead to undesirable compression artifacts, such as pumping, and undesirable loudness shifts after applying compression gain. This is caused by misaligned DRC characteristics, where the audio signal loudness is not centered within those characteristics.

[0071] 3. Improvements proposed

[0072] The proposed solution is a novel DRC processing paradigm where DRC input loudness normalization is offloaded from the encoder to the decoder. This solution eliminates the need for a first pass to measure the content's loudness before encoding begins. Instead, loudness measurement can be performed in parallel with encoding. Updated loudness values ​​can be written to the audio stream, which the decoder will use for DRC input loudness normalization. For live recording to a file, the loudness values ​​for the entire recording can be written to the MP4 file header at the end of recording, and this value can then be used for normalization in the decoder. Furthermore, there is no need to rewrite or re-encode the file to achieve the desired normalization of the DRC input to properly adjust the DRC characteristics in the decoder.

[0073] For backward compatibility, the proposed bitstream supports an efficient syntax to serve both traditional and modern decoders. This mode allows loudness normalization to be applied to the DRC chain in the encoder as needed by traditional systems. The encoder normalization gain is transmitted in the proposed bitstream syntax and compensated for when loudness normalization is applied in the modern decoder.

[0074] Sections 4 and 5 of this document contain the draft texts of the amendments to ISO / IEC 23003-4[1] and ISO / IEC 23091-3[2]. Changes and additions related to the most recently published standards are highlighted in gray to simplify the review process. The highlighting should be removed when merging the text.

[0075] 4. Proposed revisions to ISO / IEC 23003-4

[0076] In 6.1.1, replace the second to last paragraph with:

[0077] `uniDrcConfig()` contains all blocks except for the `loudnessInfo()` block bound to `loudnessInfoSet()`. The last part of the `uniDrcConfig()` payload may include future extended payloads. If the received `uniDrcConfigExtType` value is not equal to `UNIDRCCONFEXT_TERM`, the DRC tool parser should read and discard the extended payload bits (otherBit). Similarly, the last part of the `loudnessInfoSet()` payload may include future extended payloads. If the received `loudnessInfoSetExtType` value is not equal to `UNIDRCLOUDEXT_TERM`, the DRC tool parser should read and discard the extended payload bits (otherBit). Each extended payload type in `uniDrcConfig()` or `LoudnessInfoSet()` must appear more than once in the bitstream unless otherwise specified. If both payloads are present, extended payloads of type `UNIDRCCONEXT_V1` or `UNIDRCCONEXT_V2` should precede extended payloads of type `UNIDRCCONEXT_PARAM_DRC` in the bitstream. For ISO / IEC 14496-12, the configuration of extended payload is provided according to Table 76.

[0078] Replace the following paragraph in section 6.4.6:

[0079] The syntax for the intra-stream drcCoefficient is given in Tables 65, 67, and 68. The syntax for the corresponding blocks in ISO / IEC 14496-12 (ISO Basic Media File Format) is shown in Tables 66 and 69. These corresponding blocks carry essentially the same information. Except for drcLocation, the values ​​contained in both blocks are encoded in the same way.

[0080] Inventor:

[0081] The syntax for the intra-stream drcCoefficient is given in Tables 65, 67, and 68. The syntax for the corresponding blocks in ISO / IEC 14496-12 (ISO Basic Media File Format) is shown in Tables 66 and 69. These corresponding blocks carry essentially the same information. Except for drcLocation, the values ​​contained in both blocks are encoded in the same way.

[0082] The drcCoefficientsUniDrc() payload (see Table 69) of ISO / IEC 14496-12 with version=2 and characteristicV1Override=1 carries essentially the same information as the extended UNIDRCCONFEXT_V2. The corresponding bitstream fields are encoded in the same manner as specified in Table A.10.

[0083] Replace the following paragraph in section 6.4.6:

[0084] In the `drcCoefficientsUniDrcV1()` payload, custom DRC characteristics can be defined to support more flexible gain modifications. These parametric characteristics can be used to describe the encoder-side DRC and target characteristics. If a target characteristic is defined, it should be applied after inverting the encoder-side characteristic. Therefore, for the most general purpose, the encoder-side characteristic should be reversible, i.e., have a negative or positive slope across the entire gain range. Furthermore, if the target characteristic has a portion with constant gain, the encoder-side characteristic can also have constant gain in those portions. The calculation of the parametric characteristics is shown in Tables 19 and 20. More details can be found in E.4.

[0085] Inventor:

[0086] In the `drcCoefficientsUniDrcV1()` payload, custom DRC characteristics can be defined to support more flexible gain modifications. These parametric characteristics can be used to describe the encoder-side DRC and target characteristics. If a target characteristic is defined, it should be applied after inverting the encoder-side characteristic. Therefore, for the most general purpose, the encoder-side characteristic should be reversible, i.e., have a negative or positive slope across the entire gain range. Furthermore, if the target characteristic has a portion with constant gain, the encoder-side characteristic can also have constant gain in those portions. The calculation of the parametric characteristics is shown in Tables 19 and 20. More details can be found in E.4.

[0087] If encoder-side characteristics are provided in the bitstream, linear gain interpolation is recommended. ISO / IEC 23091-3 (CICP) encoder characteristics in the range of 65 to 70 are supported only by drcCoefficientsUniDrcV1() as part of the UNIDRCCONFEXT_V2 extension. These characteristics should not be used otherwise. When CICP characteristics in this range are inverted in the decoder, loudness normalization should be applied after inversion based on available loudness metadata and encoder normalized gain (if applicable). The pseudocode in Table E.3 illustrates this, where the output of the inverted characteristics is calculated using offsets based on the values ​​of sourceLoudness and encDrcNormGainDb. sourceLoudness is the synthetic DRC input loudness at the encoder before any normalization. The value of sourceLoudness is obtained from the DRC loudness metadata. encDrcNormGainDb is the signal gain applied to the encoder DRC input, in dB. When characteristicV1Override == 1, this gain value is available in the UNIDRCCONFEXT_V2 extended payload to support legacy devices (see also E.4).

[0088] The value of drcInputLoudnessTarget is the target input loudness of the DRC characteristic applied to generate DRC gain in the decoder. The drcCoefficientsUniDrc() payload of the underlying media file format ISO / IEC 14496-12 only supports CICP characteristics 65 to 70 when version >= 2 (see also Table 69).

[0089] Replace Table 17 with:

[0090] Table 17—DRC gain samples and the transformation of associated slopes from dB to the linear domain (slopeIsNegative == 1 if the source DRC characteristic has a negative slope).

[0091]

[0092]

[0093] Replace Table A.10 with:

[0094] Table A.10 — Encoding of top-level fields in uniDrcConfig(), uniDrcConfigExtension(), and loudnessInfoSet()

[0095]

[0096]

[0097] Replace Table A.16 with:

[0098] Table A.16 — Encoding of metadata in drcCoefficientsBasic(), drcCoefficientsUniDrc(), and drcCoefficientsUniDrcV1()

[0099]

[0100]

[0101]

[0102] Replace the title of Table A.22 with:

[0103] Table A.22 — Encoding of the drcCharacteristic and overrideCicpCharacteristic fields

[0104] Add a new table after Table A.22:

[0105] Table A.23—Code for bsEncDrcNormGain

[0106]

[0107] Replace Table 69 with:

[0108] Table 69—Syntax of drcCoefficientsUniDrc() payload in ISO / IEC 14496-12

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115] Replace Table 75 with:

[0116] Table 75 — Syntax of uniDrcConfigExtension() Payload

[0117]

[0118]

[0119]

[0120] Replace Table 101 with:

[0121] Table 101—Reservation Types

[0122]

[0123]

[0124] Replace Table A.12 with:

[0125] Table A.1 — UniDrc Configuration Extension Types

[0126]

[0127] Add the following paragraph before section D.2.8:

[0128] DRC source characteristics in the range of 65 to 70 specified in ISO / IEC 23091-3 (CICP) are supported only in the drcCoefficientsUniDrcV1() payload as part of the UNIDRCCONFEXT_v2 extension. They are useful for real-time content where the overall loudness of the audio signal is unknown beforehand and two-pass encoding is not possible. For such applications, loudness normalization of the DRC input signal can be performed at the decoder rather than the encoder. This "delayed" normalization provides the opportunity to obtain a loudness measurement at the encoder, which can then be used for normalization in the decoder. Based on the characteristicV1Override flag, if cleared, the UNIDRCCONFEXT_V2 extension payload contains all the information of the UNIDRCCONFEXT_V1 extension payload enhanced by supporting DRC characteristics 65 to 70; alternatively, if the flag is set, it overrides the source characteristics of the drcCoefficientsUniDrcV1() payload in the UNIDRCCONFEXT_V1 extension. The latter behavior is useful in deployments with some decoders that support the UNIDRCCONFEXT_V2 extension and others that only support the UNIDRCCONFEXT_V1 extension. While those traditional decoders can still apply DRC because the drcCoefficientsUniDrcV1() payload can be provided, they cannot take advantage of the more flexible DRC input loudness normalization with characteristics 65 to 70.

[0129] In non-backward compatible mode (characteristicV1Override == 0), the bitstream should not contain the UNIDRCCONFEXT_V1 extension, and there should be no version 1 drcCoefficientsUniDrc() payload at the file format level. In backward compatible mode (characteristicV1Override == 1), the bitstream should contain a UNIDRCCONFEXT_V1 extension that matches the pattern of the overlay information present in the UNIDRCCONFEXT_V2 extension. At the file format level, for each version 2 drcCoefficientsUniDrc() payload, there should be a version 1 drcCoefficientsUniDrc() payload that matches the pattern of the overlay information.

[0130] Replace the following paragraph in section E.4:

[0131] Calculate the inversion of transmitter characteristics 1 through 6 according to Tables E.3 and E.4. Since the gain is always 0dB, characteristic 1 does not have a useful inversion.

[0132] Inventor:

[0133] The inversion of transmitter characteristics 1 to 6 and 65 to 70 (see ISO / IEC 23091-3) is calculated according to Tables E.3 and E.4. Since the gain is always 0 dB, characteristics 1 and 65 do not have useful inversion. Section D.1.1 explains the DRC input loudness normalization in the encoder to match the DRC characteristics. This encoder-side normalization requires known content loudness, but this may not be known, for example, for real-time content. If the content loudness is unknown at the start of encoding, loudness normalization in the encoder can be omitted when using DRC characteristics with an index range of 65 to 70. These characteristics will cause the DRC decoder to apply loudness normalization, thus eliminating the need for normalization at the encoder. There are two options for applying DRC input loudness normalization in the decoder. The first option is not backward compatible. It indicates this by setting characteristicV1Override=0 and sending a UNIDRCCONFEXT_V2 extension with a drcCoefficientsUniDrcV1() payload, which includes one or more drcCharateristic values ​​in the range of 65 to 70. The encoder should not apply DRC input normalization. The second option is backward compatible. It indicates this by setting characteristicV1Override=1, sending a UNIDRCCONFEXT_V1 extension with a drcCoefficientsUniDrcV1() payload that should not include drcCharateristic values ​​in the range of 65 to 70, and sending a UNIDRCCONFEXT_V2 extension payload with an overlay characteristic index (including indices in the range of 65 to 70) and the encoder normalized gain value bsEncNormGainDb. DRC input normalization is typically applied in encoders for backward compatibility with legacy decoders.

[0134] For DRC with target characteristics, it is recommended to use linear DRC gain coding (gainInterpolationType=1) and ignore slope information. If the gain coding is spline interpolation, then use linear interpolation in the decoder.

[0135] Replace Table E.3 with:

[0136] Table E.3—Calculation of DRC characteristics 1 to 6 and 65 to 70 for inverting encoders

[0137]

[0138]

[0139] Replace Table E.4 with:

[0140] Table E.4—Parameters for DRC characteristics with index ranges of 1 to 6 and 65 to 70

[0141]

[0142] Replace Table I.2 with:

[0143] Table I.2 — In-stream payload processing supported by each configuration file

[0144]

[0145]

[0146] Replace Table I.3 with:

[0147] Table I.3—Support for MPEG-4 file format frame processing for each configuration file

[0148]

[0149]

[0150] 5. Proposed revisions to ISO / IEC 23091-3

[0151] Replace the following content in section 6.12:

[0152] All DRC features are defined based on the DRC input level with loudness normalized to -31LKFS.

[0153] Inventor:

[0154] Features with an index range of 1 to 11 are defined using DRC input levels with loudness normalized to -31LKFS. For features with an index range of 65 to 70, loudness normalization should not be used for the DRC input.

[0155] Replace the following content in section 6.12:

[0156] Table 6—Parameters with DRC characteristic index values ​​from 1 to 6

[0157]

[0158] Inventor:

[0159] Table 6—Parameters of DRC characteristics with index ranges of 1 to 6 and 65 to 70

[0160]

[0161]

[0162] 6 Conclusions

[0163] This document proposes improvements to the DRC standard for real-time coding. This includes proposed draft amendments to ISO / IEC 23003-4 and ISO / IEC 23091-3. The submitter requests that the draft amendments to both standards be published using the provided text.

[0164] 7 References

[0165] [1] ISO / IEC 23003-4:2020 Information technology—MPEG audio technology—Part 4: Dynamic range control.

[0166] [2] ISO / IEC 23091-3:2018 Information technology—Coding of independent code points—Part 3: Audio.

Claims

1. An audio decoder device, comprising: processor; as well as A memory storing instructions that configure the processor to acquire a bitstream, the bitstream comprising: The encoded version of the audio signal; A first dynamic range control (DRC) gain sequence is determined by an encoding process that applies the audio signal to a first DRC characteristic. Loudness normalized gain, which is applied by the encoding process when determining the first DRC gain sequence; An index of the first DRC feature, wherein the index identifies or points to the first DRC feature; and Multiple instances of integrated loudness updates over time.

2. The audio decoder apparatus of claim 1, wherein, in response to the index having a first value, the processor performs loudness normalization when applying DRC to the audio signal.

3. The audio decoder device according to claim 1, The bitstream indicates that the processor performs loudness normalization after applying the inverse DRC feature to the first DRC gain sequence, by using the loudness normalization gain in the bitstream to compensate for or undo the loudness normalization gain applied by the encoding process when determining the first DRC gain sequence.

4. The audio decoder apparatus according to any one of claims 1 to 3, wherein the memory stores instructions therein configuring the processor to: The loudness sequence is recovered by applying the first DRC gain sequence to the inverse characteristic of the first DRC characteristic; Perform loudness normalization on the recovered loudness sequence; A second DRC gain sequence is generated by applying the recovered loudness sequence to a second DRC characteristic; and The second DRC gain sequence is applied to the audio signal.

5. The audio decoder apparatus of claim 4, wherein the loudness normalization gain is in dB, and performing loudness normalization includes combining the loudness normalization gain with the recovered loudness sequence and an instance of the integrated loudness update.

6. The audio decoder apparatus of claim 4, wherein performing the loudness normalization comprises shifting the second DRC feature along its input axis by an amount based on the loudness normalization gain and an instance of the integrated loudness update.

7. The audio decoder apparatus of claim 4, wherein for each instance of the integrated loudness update, the processor calculates an update to the normalized gain as the difference between the DRC input loudness target and the instance of the integrated loudness update, and performs the loudness normalization on the recovered loudness sequence by adding the normalized gain to the recovered loudness sequence to produce a normalized loudness sequence, and wherein the processor produces the second DRC gain sequence by applying the normalized loudness sequence to the second DRC feature to produce the second DRC gain sequence.

8. The audio decoder apparatus of claim 1, wherein adjacent instances of the integrated loudness update are spaced one to ten seconds apart.

9. The audio decoder apparatus of claim 1, wherein the integrated loudness update represents the running average integrated loudness of the audio signal.

10. The audio decoder apparatus of claim 1, wherein the processor is configured to: Extract the index of the first DRC feature from the bit stream and use the extracted index to obtain the inverse feature of the first DRC feature; The loudness sequence is recovered by applying the first DRC gain sequence to the inverse characteristic of the first DRC characteristic; If the index has a first predefined value, then for each instance of the integrated loudness update, a normalized gain update is computed as i) the difference between the DRC input loudness target and ii) the sum of the instance of the integrated loudness update and the loudness normalized gain, and the normalized gain update is added to the recovered loudness sequence to produce a normalized loudness sequence. A second DRC gain sequence is generated by applying the normalized loudness sequence to the second DRC characteristic; as well as The second DRC gain sequence is applied to the audio signal.

11. The audio decoder apparatus of claim 10, wherein the processor is configured to generate the second DRC gain sequence by applying the recovered loudness sequence without loudness normalization to the second DRC feature if the index has a second predefined value.

12. An audio decoder device, comprising: processor; as well as A memory storing instructions that configure the processor to acquire a bitstream, the bitstream comprising: The encoded version of the audio signal; A first dynamic range control (DRC) gain sequence is determined by an encoding process that applies the audio signal to a first DRC characteristic. An index of the first DRC feature, wherein the index identifies or points to the first DRC feature; Multiple instances of integrated loudness updates over time; and A flag, wherein when the flag has a first value, the bitstream includes the encoder-side loudness normalization gain, or when the flag has a second value, the bitstream does not include the encoder-side loudness normalization gain.

13. The audio decoder apparatus of claim 12, wherein in response to the flag having the first value, the processor applies DRC to the audio signal using a second DRC feature, and performs loudness normalization using i) the encoder-side loudness normalization gain and ii) the plurality of instances of integrated loudness updates.

14. The audio decoder apparatus of claim 12, wherein in response to i) the flag having the second value and ii) when the index has the first value, the processor applies DRC to the audio signal using a second DRC feature and uses the plurality of instances of integrated loudness update, but does not use the encoder-side loudness normalization gain for loudness normalization.

15. The audio decoder apparatus of claim 14, wherein in response to the index being a second value different from the first value, the processor uses the second DRC feature to apply DRC to the audio signal without loudness normalization.

16. An audio decoder device, comprising: processor; as well as A memory storing instructions that configure the processor to acquire a bitstream, the bitstream comprising: The encoded version of the audio signal; A first dynamic range control (DRC) gain sequence is determined by an encoding process that applies the audio signal to a first DRC characteristic. An index of the first DRC feature, wherein the index identifies or points to the first DRC feature; Multiple instances of integrated loudness updates over time; and The processor replaces some or all of the traditional DRC feature index values ​​of the traditional extended payload in the bitstream with DRC feature index values ​​from the new extended payload contained in the bitstream when the flag has a first value.

17. An audio decoder device, comprising: processor; as well as A memory storing instructions that configure the processor to acquire a bitstream, the bitstream comprising: The encoded version of the audio signal; A first dynamic range control (DRC) gain sequence is determined by an encoding process that applies the audio signal to a first DRC characteristic. An index of the first DRC feature, wherein the index identifies or points to the first DRC feature; and Multiple instances of integrated loudness updates over time The bitstream includes encoding-side DRC normalization gain, which the processor compensates for when applying decoding-side DRC loudness normalization.

18. A digital audio method, comprising: A bitstream is obtained, the bitstream including an encoded version of an audio signal, a first dynamic range control DRC gain sequence determined by an encoding process of applying the audio signal to a first DRC feature, an index of the first DRC feature, wherein the index identifies or points to the first DRC feature, and multiple instances of integrated loudness updates over time. Use the index to obtain the inverse DRC property; After applying the inverse DRC characteristic to the first DRC gain sequence, loudness normalization is performed to produce a normalized loudness sequence. The normalized loudness sequence is applied to the second DRC characteristic to generate a second DRC gain sequence; as well as The second DRC gain sequence is applied to the audio signal to produce compressed audio.

19. The method of claim 18, wherein the bitstream includes a loudness-normalized gain applied by the encoder when determining the first DRC gain sequence by applying the audio signal to the first DRC characteristic. The bitstream instructs the processor of the audio decoder device to perform loudness normalization by using the loudness normalization gain in the bitstream to compensate for or undo the loudness normalization gain applied by the encoder when determining the first DRC gain sequence.

20. The method of claim 18, wherein the bitstream includes a flag, and when the flag has a first value, the encoding process applies the audio signal to the first DRC characteristic and performs loudness normalization to determine the first DRC gain sequence.

21. The method of claim 20, wherein when the flag has a second value, the first DRC gain sequence is determined by the encoding process applying the audio signal to the first DRC characteristic without loudness normalization.

22. The method according to any one of claims 18 to 21, wherein performing loudness normalization comprises: The normalized loudness sequence is adjusted, and then the adjusted loudness sequence is applied to the second DRC characteristic.

23. An audio decoder device, comprising: processor; as well as A memory storing instructions that configure the processor to obtain a bitstream, the bitstream including an encoded version of an audio signal, an instantaneous loudness sequence, and an integrated loudness value; The normalized gain is calculated by combining the DRC input loudness target with the integrated loudness value extracted from the bitstream; The normalized gain is used to adjust the instantaneous loudness sequence extracted from the bitstream to produce a normalized instantaneous loudness sequence; as well as A DRC gain sequence is generated by applying the normalized instantaneous loudness sequence to the DRC characteristics; as well as DRC is performed on the audio signal by applying the DRC gain sequence to the audio signal.

24. The audio decoder apparatus of claim 23, wherein the instantaneous loudness sequence in the bitstream has not been loudness normalized.

25. The audio decoder apparatus according to any one of claims 23 to 24, wherein the integrated loudness value is one of a plurality of instances of integrated loudness updates contained in the bitstream, adjacent instances being spaced one to ten seconds apart, wherein the integrated loudness update represents the running average integrated loudness of the audio signal.

26. The audio decoder apparatus according to any one of claims 23 to 24, wherein the bitstream is a file in which the integrated loudness values ​​have been written together with the instantaneous loudness sequence and the encoded version of the audio signal.

27. A digital audio method, comprising: Obtain a bitstream, which includes an encoded version of an audio signal, an instantaneous loudness sequence, and an integrated loudness value; The normalized gain is calculated by combining the DRC input loudness target with the integrated loudness value from the bitstream; The instantaneous loudness sequence from the bitstream is adjusted using loudness normalization gain to produce a normalized instantaneous loudness sequence; A DRC gain sequence is generated by applying the normalized instantaneous loudness sequence to the DRC characteristics; as well as DRC is performed on the audio signal by applying the DRC gain sequence to the audio signal.

Citation Information

Patent Citations

  • Dynamic range and peak control in audio using nonlinear filters

    US10109288B2

  • Metadata for loudness and dynamic range control

    CN105103222A

  • Encoded audio metadata-based loudness equalization and dynamic equalization during DRC

    CN107925391A