Dynamic equalization during loudness equalization and DRC based on encoded audio metadata

By using metadata-based loudness equalization and dynamic equalization techniques, the problem of low-frequency range loss in audio playback is solved, achieving consistency and high-quality playback of audio content at different playback levels, while reducing device complexity and latency.

CN114070217BActive Publication Date: 2026-03-20APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2016-09-26
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the issue of low-frequency range loss in audio playback, resulting in inconsistent listening experiences at different playback levels and affecting the tonal characteristics of audio content.

Method used

By employing metadata-based loudness equalization and dynamic equalization techniques, time-varying filters are used to compensate for spectral distortion. Time-varying filtering is performed based on the time-varying level of the audio content's spectral bands to adjust the audio signal before playback and achieve spectral shaping.

Benefits of technology

It improves the consistency and quality of audio playback, reduces the complexity and latency of playback devices, and enhances the listener's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114070217B_ABST
    Figure CN114070217B_ABST
Patent Text Reader

Abstract

The present disclosure is based on dynamic equalization during loudness equalization and DRC of encoded audio metadata. The invention provides for dynamic loudness equalization of received audio content in a playback system using metadata including a momentary loudness value for the audio content. A playback level is derived from a user volume setting of the playback system and compared to a mix level assigned to the audio content. A parameter is calculated that defines an equalization filter to filter the audio content before driving a loudspeaker with the filtered audio content, the parameter being based on the momentary loudness value and the comparison of the playback level to the assigned mix level. Other embodiments are described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of PCT International Application No. 201680050105.5, filed September 26, 2016, entitled "Dynamic Equalization During Loudness Equalization and DRC Based on Encoded Audio Metadata," which entered the Chinese national phase as of September 26, 2016.

[0002] This patent application claims the benefit of earlier filing date of U.S. Provisional Patent Application 62 / 235,293, filed September 30, 2015. TECHNICAL FIELD

[0003] Embodiments of the invention relate to playback-side digital audio signal processing of digital video content associated with metadata to improve the experience of a listener. Other embodiments are also described. BACKGROUND

[0004] Audio content such as film music or soundtracks is typically produced with the assumption of a certain playback level (e.g. an "overall gain" that should be applied to the audio signal during playback, between its initial or decoded form and its conversion to sound by a loudspeaker, in order to obtain a sound pressure level at the same listener position as intended by the audio content producer). If a different playback level is used, the content not only sounds louder or softer, but can also appear to have different tonal characteristics. An effect known from psychoacoustics is a non-linear increase of low frequency loudness perception as a function of playback level. This effect can be quantified by equal-perceived loudness contours as a function of playback level and signal characteristics and by a measure of perceived loudness. Typically, a partial loss of low frequency content is reported when the content is played back at a lower level than intended by the producer, compared to other frequencies. In the past, loudness equalization was done by an adaptive filter that amplified the low frequency range depending on the playback volume setting. Many older audio receivers have a "loudness" button that works in this way. SUMMARY

[0005] Several schemes for metadata-based loudness equalization (EQ) are described below. Some of the schemes can have one or more of the following advantages: e.g. reduced playback-side complexity, less delay and higher quality. Some of the quality improvements can be due to offline processing at the encoding side, which is not limited by real-time processing constraints and low delay requirements in playback devices. The metadata-based approach described herein can also seamlessly integrate into the existing MPEG-D DRC standard ISO / IEC, "Information technology - MPEG audio technologies - Part 4: Dynamic range control," ISO / IEC 23003-4:2015, and work together with dynamic range control.

[0006] A method is also described to provide dynamic EQ within the DRC process. It enables similar EQ as multi-band DRC but with a lower number of bands or just a single band DRC. The dynamic EQ can be controlled by metadata and can be integrated into the common MPEG-D DRC standard.

[0007] The above summary does not include an exhaustive list of all aspects of the present application. It is contemplated that the present application includes all systems and methods that can be practiced with the claimed subject matter as disclosed in the various aspects set forth above and as disclosed in the detailed description below and particularly pointed out in the claims filed with the present patent application. Such aspects have particular significance to the present application. BRIEF DESCRIPTION OF DRAWINGS

[0008] Embodiments of the present application are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements. It should be noted that references made to "an" or "one" embodiment in this disclosure are not necessarily to the same embodiment, and that such references mean that at least one. In addition, to be clear and concise, a certain drawing can be used to show features of more than one embodiment of the present application, and not all elements in such a drawing can be required for a certain embodiment.

[0009] Figure 1 is a block diagram of a decoding side loudness equalizer based on instantaneous loudness values extracted from received metadata of audio content.

[0010] Figure 2 is a block diagram of a generating or encoding side system for generating metadata including instantaneous loudness values.

[0011] Figure 3 shows several example DRC characteristics that can be used in the encoding side to calculate DRC gain values and in inverse form in the decoding side.

[0012] Figure 4 is a schematic diagram showing how to generate instantaneous loudness values in the decoding side using inverse DRC characteristics for adjusting a loudness equalization filter.

[0013] Figure 5 shows a decoding side where forward audio content is applied dynamic range compression and loudness equalization.

[0014] Figure 6 shows a decoding side where forward audio content is applied dynamic range compression and dynamic equalization.

[0015] Figure 7 is a block diagram of another system for loudness equalization in the decoding side. DETAILED DESCRIPTION

[0016] Several embodiments of the invention will now be explained with reference to the accompanying drawings. Where the shape, relative position, and other aspects of the components described in the embodiments are not explicitly defined, the scope of the invention is not limited to the components shown, which are for illustrative purposes only. Furthermore, while many details have been set forth, it should be understood that some embodiments of the invention can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0017] The conventional "loudness" button mechanism described in the background section above overlooks an important issue: the amount of low-frequency range loss reported by the listener depends on the acoustic level of that frequency range at the listener's location, which in turn depends on the audio content. An embodiment of the present invention is a loudness equalization scheme that takes into account the time-varying level of the spectral band of the audio content to control the passage of the audio signal (before driving the speaker) through a time-varying filter. This time-varying filter (spectral shaping filter, also referred to herein as an equalization filter) is designed to compensate for spectral distortion that will occur as a function of playback level and frequency band due to nonlinear loudness perception.

[0018] Figure 1 The diagram provided illustrates the concept of a loudness equalizer that works based on metadata, which is related to... Figure 1 The diagram above is associated with the audio content being played back. This diagram (and all other diagrams in this document) indicates a digital signal processing (DSP) control or DSP logic hardware unit (e.g., a processor that executes instructions stored in a machine-readable medium such as local storage or memory in a home audio system, consumer electronics speaker device, or in-vehicle audio system), and also indicates a decoding and playback system that receives the audio content, such as a desktop computer, home audio entertainment system, set-top box, laptop computer, tablet computer, smartphone, or other electronic audio playback system, in which the resulting digital audio output signal is converted into analog form and then fed to an audio power amplifier that drives speakers (e.g., speakers, headphones). Audio content initially received, for example, via an Internet stream or Internet download, may have already been encoded and multiplexed with its metadata into a bitstream by the time it reaches the processing shown in the diagram, which has already been decapsulated and decoded in the playback system.

[0019] The metadata 2 comprises static metadata, including a mixing level of the audio content and optionally a program loudness, e.g. as a single value for the complete content (herein also referred to as audio program or audio asset) respectively. The mixing level can be measured during production (or at the encoding stage) by means of the following well-established standard. The program loudness value can be measured using a loudness model, such as the model defined in ITU "Algorithms to measure audio programme loudness and true-peak audio level," ITU-R BS.1770-3. In addition, instantaneous loudness values (e.g. acoustic level, sound pressure level, SPL, etc.) are conveyed as dynamic metadata via a metadata stream, wherein several instantaneous loudness values are received successively over time, which describe the sound pressure level (SPL) of the audio content at the intended position of the listener for each frame or block (synchronized with the frames of the audio content signal). In other words, the instantaneous loudness changes within a time sequence of such frames or blocks defining the audio content. The metadata can be transmitted to the illustrated playback or decoding side, shown as a decoding and playback system, via e.g. Internet download or via Internet streaming together with the audio content (denoted as "audio input" in the figure). At the decoding or playback side, no additional delay is caused, since the instantaneous loudness values are in the metadata, hence a loudness estimation process at the playback side is not necessary. Improved smoothness, reduced decoder complexity and no additional delay are all advantages of this approach over prior art techniques for loudness equalization, which operate exclusively at the playback side without using metadata.

[0020] At the playback side, the user's volume setting (for manual control of the sound volume from the loudspeaker or headphones during playback) is input to a volume control block 4. The volume control block 4 then generates (e.g. computes, possibly including a table lookup) the appropriate gain value (e.g. full-band scaling factor), which is applied to the digital audio output signal ( "audio output" in the figure). It also derives the playback level based on the user volume setting and based on its stored or predetermined knowledge of the playback system's level transmission characteristics (sensitivity). The latter describes how the given audio output signal is rendered as sound with the resulting sound pressure level in the listener's area (note that this sensitivity can also depend on factors such as the user volume setting).

[0021] The filter adjustment block 7 takes the metadata 2 and the calculated difference between the static mix level (indicated in the metadata) and the playback level (e.g. as a subtraction between the two dB values, also referred to as comparing the mix and playback levels) and based on this difference, generates filter parameters that control (e.g. define) the equalization (EQ) filter 5. The filter adjustment block 7 can first determine whether the playback level is higher or lower than the mix level. (If the metadata does not provide a mix level, an average mix level can be employed (e.g. typically used in sound program or audio recording production environments)). If the playback level is lower, then the low frequency range, optionally the high frequency range, needs to be boosted to some extent depending on the instantaneous loudness reported for the audio content (metadata 2). Similarly, if the playback level is higher, then these spectral ranges need to be attenuated to some extent. The EQ filter 5 is configured to do so and can be updated for each frame (e.g. for each frame of the digital audio content (audio input) that has an instantaneous loudness value associated with it in the metadata, or by skipping some frames, so that the EQ filter 5 does not need to be updated for each frame of the audio content).

[0022] Note that in case the playback level is lower than the mix level, the amount of boost imparted by the EQ filter 5 is greater, the lower the playback level is compared to the mix level, the lower the instantaneous loudness value is. This is because the human non-linear loudness perception increases relative to level at low sound pressure levels. Also, in one embodiment, in case the playback level is not sufficiently different from the mix level, no spectral shaping (e.g. its response should be flat at 0 dB) is required by the EQ filter 5.

[0023] Typically, it is advantageous to split the audio spectrum into several frequency bands and estimate the loudness in each of these bands individually. For the case of loudness equalization addressed in particular herein, a frequency band for low frequencies and another (non-overlapping) frequency band for high frequencies can be defined for the loudness measurement at the encoding side (rendered as a sequence of pairs of instantaneous loudness values in the metadata). This is done as an attempt to model the human auditory perception in those frequency ranges. Alternatively, instantaneous loudness values can be provided for only a single frequency band, e.g. low frequencies below 200 Hz. These loudness values are then adapted to control the EQ filter 5 in the manner conceptually described above.

[0024] In one embodiment, the information required to control the EQ filter 5 includes (I) the instantaneous SPL in a particular audio frequency band (spectral range) for the audio content at the mix (production) and (II) at the playback. In the following, (I) is referred to as L range,mix (t) and (II) is referred to as L range,playback (t). Using such inputs, the required boost gain or cut gain in a particular frequency band can then be computed using conventional means.

[0025] Given this audio content, the instantaneous loudness level can be estimated in the audio frequency band at the production or encoding side, but only the absolute level during playback can be determined when the sensitivity of the playback system is known. For a playback system, the sensitivity ΔL playback describes the measured difference between the acoustic level [sound pressure level] L playback and the electrical audio signal level [dBFS] of the content that causes this level. For a production system, the sensitivity can be defined as the difference between the electrical audio signal level (L content ) of the content and the resulting measured SPL, e.g. the mix level in SPL. ΔL mixing The sensitivity of a mixing system can be included as a static value in the metadata. Alternatively, it can be estimated in the playback system by calculating the difference between the mix level (e.g. the average SPL measured in a mix room) and the average loudness level (both values can be conveyed as metadata). The average loudness level can be calculated, for example, by the methods described in ITU-R BS.1770-3. This estimation is referred to as program loudness - see Figure 1 Typically, the sensitivity of a production / mixing system is constant. However, for the playback system, it can vary when the user adjusts the volume, e.g. by turning the volume knob of the device.

[0026] If the instantaneous loudness value [SPL] cannot be calculated or measured in the mix room, then it can be calculated at the playback side, based on the absolute content level L range,content in [dBFS] (t) where (t) indicates that it changes over time due to level fluctuations of the content (audio input). With this estimation, the SPL in the spectral range can be calculated for the mixing side and the playback side:

[0027] L range,mixing (t) = L range,content (t) - deltaL mixing

[0028] and

[0029] L range,playback (t) = L range,content (t) + deltaL playback

[0030] Alternatively, the average level difference ΔL mixing between the production and playback side can be calculated directly based on the average mix level L playback and the playback level L acoustic :

[0031] ΔL acoustic = L playback - L mixing

[0032] Based on this result, the instantaneous SPL at the playback side is:

[0033] L range,playback (t) = AL acoustic + L range,mix (t)

[0034] The human perception of loudness in certain spectral ranges decreases non-linearly at lower sound pressure levels SPL in the low frequency range. There exists a conventional perception of loudness that can be measured in the laboratory environment for low and medium range in mixing and in playback situations.

[0035] The amount of such boost gain can be calculated using conventional techniques based on measurements of the various test signals of various publications depending on the psychoacoustics of frequency and level. For example, see T. Holman and F. Kampmann, "Loudness Compensation: Use and Abuse", Journal of the Audio Engineering Society, July / August 1978, Vol. 26, No. 2 / 8. A common representation of the data is in the form of an isophonic curve or graph showing the loudness growth versus level. With such psychoacoustic data, the amount of boost gain can be easily calculated (by programming the filter adjustment block 7) as a function of the instantaneous loudness value and based on the playback and mixing levels as described above. Based on this boost gain value, and the frequency band for which a frequency boost should be applied, the parameters of the digital filter elements, which are part of the EQ filter 5 and will produce the appropriate boost in the frequency range of interest, can be derived.

[0036] Example EQ filter element

[0037] The following shows example cut-off and boost filter elements for the low and high frequency range, which can approximate the desired frequency response for loudness equalization. In this example, several different filter elements are connected to form a cascade as part of the Figure 1 EQ filter 5, where each element can have a cut-off or boost frequency range while leaving the rest of the audio spectrum unchanged (0 dB gain input). These are examples of low and high shaping filters, where the low and high shaping filters form a cascade as part of the EQ filter 5.

[0038] Each low shaping filter (i.e. part of the EQ filter 5) can be a first order IIR filter with real coefficients, which has the form:

[0039]

[0040] The low-frequency shelving filter can have fixed coefficients a1 depending on the desired corner frequency. The filter parameters a1 can be dynamically computed based on a linear gain L boost or L as defined above boost The filter parameters b1 are dynamically computed as follows:

[0041]

[0042] The low-frequency shelving filter can have fixed coefficients a1 depending on the desired corner frequency. The filter parameters a1 can be dynamically computed based on a linear gain L boost The filter parameters a1 are dynamically computed as follows:

[0043]

[0044] Each high-frequency shaping filter can be a second-order IIR filter with real coefficients, having the form:

[0045]

[0046] The corner frequency of the filter can depend on the audio sampling rate and a normalized corner frequency

[0047] f c = f c,norm f s

[0048] Each high-frequency shelving filter can have fixed coefficients, except for b1. The fixed filter coefficients depend on the corner frequency exponent, as well as the pole / zero radius parameters.

[0049] r = 0.45

[0050] a1 = -2rcos(2πf c,norm )

[0051] a2 = r 2

[0052] b2 = a2

[0053] The filter parameters b1 can be dynamically computed based on a linear gain L boost or L boost The filter parameters b1 are dynamically computed as follows:

[0054]

[0055] The high-frequency shelving filter can have the same coefficients, except that the a coefficients are computed in the same way as the b coefficients for the shelving filter, and the b coefficients are computed in the same way as the a coefficients for the shelving filter.

[0056] b1 = -2rcos(2πf c,norm )

[0057] b2 = r2

[0058] a2 = b2

[0059] The boosting gain L can be based on boost The filter parameter a1 is dynamically calculated:

[0060]

[0061] Figure 2 An example production / encoding side system is shown in Fig. 1 for producing metadata including instantaneous loudness (e.g. SPL) values for a given audio program. In order to improve the accuracy of the instantaneous loudness or SPL measurement performed at the encoding side, for the purpose of better downstream control of the EQ filter 5 at the playback side, the audio signal can first be processed by a band-pass filter 13 which removes all components outside the spectral range of interest (the spectral range to be modified by the EQ filter 5) before the audio signal enters the loudness measurement module 14. In this way, a more accurate estimation of the instantaneous loudness can be achieved to obtain a better perceptual quality of the loudness equalization in the spectral range. For example, the instantaneous loudness can be derived by computing the short-term energy at the output of the band-pass filter 13 by the loudness measurement module 14 and then smoothing the computed short-term energy sequence (smoothing block 16) to avoid rapid fluctuations of the instantaneous loudness values. For the case of offline processing of the audio content (as compared to live streaming), the look-ahead of the smoothing can be increased to improve the smoothing and to avoid distortions which can otherwise occur if the EQ filter 5 is adjusted too fast or not at the right time (in response to the instantaneous loudness values).

[0062] According to another embodiment of the present application, the need to include a sequence of instantaneous loudness values (computed at the production or encoding side) in the metadata is avoided by exploiting the following way (to implement a loudness EQ in the decoder side / playback system). ISO / IEC, “Information technology - MPEG audio technologies - Part 4: Dynamic range control,” ISO / IEC 23003-4:2015 defines a flexible scheme for loudness and dynamic range control (DRC). The scheme uses a sequence of gains within the metadata to convey DRC gain values to the decoder side to apply a compression effect in the decoder side by applying the DRC gain values to the decoded audio signal. Back at the encoder, these DRC gain values are typically produced by applying a DRC characteristic, e.g. the DRC characteristic shown in Fig. 2, to a smoothed instantaneous loudness estimate. Figure 3 Figure 3 ​From ISO / IEC, “Information technology–MPEG audio technologies–Part 4: Dynamic range control,” ISO / IEC 23003-4:2015. Figure 3 The DRC input level in the curve is the smoothed instantaneous loudness level.

[0063] According to embodiments of the invention, a DRC gain sequence sourced from the same metadata can be used, intended for compressing audio programs, such as those defined in MPEG-D DRC, for loudness equalization (when the same audio program is played back). Reference Figure 4 This can be accomplished by applying the inverse DRC characteristic function 20 to the DRC gain sequence (on the decoding side), which recovers the smoothed instantaneous loudness value, reinterprets it as an instantaneous SPL value, and then uses it to dynamically update the loudness EQ filter 5, as described above. The inverse function can be obtained, for example, by inverting the input and output variables of a mathematical function that is one of several DRC gain curves (e.g., Figure 3 (As shown in the diagram), the DRC gain curve is, or represents, a sequence of coded DRC gain values ​​received in the metadata, calculated using DRC characteristics already applied on the encoding side. In other words, the inverse DRC characteristic can be the inverse function of the DRC characteristic applied on the encoding side to the audio content to produce DRC gain values. This latter sequence is now applied to the "output" of a mathematical function (or as the input to the calculated inverse function of the mathematical function) to produce a corresponding loudness sequence for each DRC frame, which is treated as an instantaneous loudness level. Note that a reference level offset can be applied to adjust each such instantaneous value of the sequence, for example, where the offset is a fixed value in dB, representing a reference acoustic level, and then the offset-adjusted sequence is fed to filter adjustment block 7. Figure 4 All other aspects can be with Figure 1 The same applies, including optional transformations to linear block 22 ( Figure 1 (Not shown in the image) may require a linear block to convert the dB value calculated by volume control 4 into a linear format, and then scale or double the filtered audio content (from the decoded audio program) that appears from EQ filter 5 to reflect the current user volume setting.

[0064] According to another embodiment of the application, it is recognized that it is also useful to have separate DRC gain sequences in the metadata that are targeted for loudness EQ only. For this purpose, the MPEG-D DRC standard can be extended by additional metadata syntax that carries the information which of the several gain sequences (included in the metadata) is suitable for loudness EQ (to control the EQ filter 5) and which frequency range should be controlled. There can be several such dedicated DRC gain sequences in the metadata, each dedicated for performing loudness EQ on a different frequency range. In addition, the additional metadata can specify which of the gain sequences to be used for loudness EQ is also suitable for specific downmixing and, if applicable, dynamic range control. This embodiment is illustrated by the block diagram of Figure 5 . Figure 5 and Figure 4 The similarities between Figure 5 and Figure 4 are obvious, while the differences include: in Figure 5 , dynamic range control (DRC gain adjustment) is applied at the multiplier and is derived from a different DRC gain sequence, sequence 2 (in this case with an optional DRC gain modification block 25), while in Figure 4 , no DRC gain adjustment is applied. The EQ filter 5 (loudness EQ) is now controlled as a function of the instantaneous loudness (SPL) value derived from the DRC gain sequence 1 of unique purpose, and corrected by the DRC gain adjustment value (at the summation unit) that is simultaneously applied at the multiplier for dynamic control (which is derived from the DRC gain sequence 2).

[0065] is also in Figure 5In this case, the static offset applied to the instantaneous loudness values is not, for example, a fixed reference SPL, but a dynamic correction is made, which can be given by the output of the DRC gain modification block 25. Block 25 is optional, however, as the correction made to the instantaneous loudness values can alternatively be given directly by the DRC gain sequence 2 originating from the metadata. The DRC gain modification block 25 can be optionally included in comparison to the generation / encoding side selection and use (to calculate the DRC gain sequence originating from the metadata) in order to change the compression curve or DRC characteristic applied during playback. The DRC gain modification block 25 can be in accordance with the specification in US patent application publication 2014 / 0294200 (

[0040] -

[0045] paragraphs, generating so-called "modified" DRC gains (new DRC gain adjustment values)), which can be more suitable for this particular playback system. In either case, the instantaneous loudness sequence input to the filter adjustment block 7 is now corrected by the DRC gain value sequence, the gain values of which are also applied to the scaled audio content, for example, as shown in the figure downstream of the EQ filter 5, to perform dynamic range control. Thus, with such a technique, if the metadata already provides DRC gain values on a frame-by-frame basis, there is no need to include a separate sequence of instantaneous loudness values in the metadata (in combination with Figure 1 See above (to implement playback side loudness EQ).

[0066] In yet another embodiment, now referring to Figure 6 , Figure 4 the loudness EQ scheme of Figure 5 (also combining loudness EQ and DRC) is combined with DRC, whereby the dynamically range adjusted audio content is filtered by the EQ filter 5 to finally produce EQ filtered and dynamically range adjusted audio content. This is done, however, in a different way than in Figure 5 The differences with respect to Figure 5 include: the instantaneous loudness sequence (in the metadata) provided with a given DRC gain sequence 1 is mixed to the filter adjustment block 7 by adding a reference level (e.g. a fixed value) to the applied inverse DRC characteristic function 20; and the input of the filter adjustment block 7 is dynamically updated by adjusting the static difference between the playback levels and dynamically mixing the levels according to the simultaneously applied DRC gain. Other differences include the addition of a DRC interface, which is the format of the control parameters fed to the DRC block, such as defined in MPEG-D DRC, and the application of the DRC gain to the audio content upstream of the EQ filter 5 (as compared to downstream of the EQ filter 5 as shown in

[0067] Dynamic equalization and DRC

[0068] The above described approach provides a loudness EQ tool that can be combined with a DRC such as the one provided in MPEG-D DRC. For some applications, however, the loudness EQ tool can be too complex or can not be properly controlled, e.g. because the playback level is not known.

[0069] In many applications, dynamic range compression is implemented using a multi-band DRC. In many cases, it is also possible to "dynamic equalize" by individually controlling the compression in each DRC band. The following approach provides such a dynamic EQ approach for purposes other than just loudness EQ.

[0070] In the following, a dynamic EQ approach for a DRC is described that works in a similar way as the above described loudness EQ. The difference is that it does not take the playback level into account (e.g. as produced by the volume control block 4, see Fig. 5, 6). Instead it applies an EQ in order to compensate for the coloring effect of the dynamic range control that can partly result from the association of the psychoacoustic properties of loudness perception with level changes caused by the application of the DRC. Other useful applications of the dynamic EQ approach described herein include, e.g., band-pass filtering of low-level background sounds in the audio content to avoid a large amplification of the noise that can otherwise cause the sound to be annoying. Figure 1 、 4

[0071] The approach described in the following can be integrated into MPEG-D DRC (metadata based). But it can also be used for common real-time dynamic range control (without metadata). It can provide the benefits of dynamic EQ for single-band DRC, which was only possible in the context of a multi-band metadata based DRC process before. A conventional metadata based single-band DRC process applies the same gain to all frequency components, so that it is not possible to selectively reduce the DRC gain, e.g. only in the low frequency range.

[0072] Furthermore, the approach described in the following is not limited to the sub-band resolution of a conventional metadata based multi-band DRC process, so that it can provide a smoother spectral shaping and can have a lower computational complexity. Figure 7 An example of a DRC capability combined with dynamic EQ is shown. The EQ is indirectly dynamically controlled by the DRC gain sequence, where this aspect is similar to Figure 5 ​part, as it includes the inverse DRC characteristic function 20, the summation unit of the instantaneous loudness corrected by the DRC gain values, and the optional DRC gain modification block 25 (the DRC gain values produced by block 25 are converted into loudness dB values by the conversion to dB block 26). Here, however, the loudness values are obtained by applying the inverse DRC characteristic function 20 to the same metadata-derived DRC gain sequence that produces the DRC gain values. In the present embodiment, the EQ filter 5 is set based in part on static metadata communicated in the bitstream that determines: filter type (e.g., low frequency cut / boost, high frequency cut / boost), filter strength, and adjustment frequency range. It is noted here that in the alternative embodiment of the equalizer in 5, static filter configuration information in the metadata is not required. Figure 1 , 4 , 5. In the alternative embodiment of the equalizer in 5, static filter configuration information in the metadata is not required.

[0073] In the alternative embodiment of the equalizer in 5, static filter configuration information in the metadata is not required. Figure 7 In the alternative embodiment of the equalizer in 5, static filter configuration information in the metadata is not required.

[0074] The following statements are now made regarding the present invention. An article of manufacture includes a machine readable medium having stored therein instructions which, when executed by a processor of an audio playback system, perform dynamic audio equalization while applying dynamic range control as follows. Audio content is received, and metadata for the audio content is also received, where the metadata includes a plurality of dynamic range control, DRC, gain values that have been computed for the audio content. The plurality of DRC gain values received in the metadata are applied to inverse DRC characteristics to compute a plurality of instantaneous loudness values for the audio content. A plurality of dynamic parameters defining an equalization filter are computed, where the dynamic parameters are computed based on the computed plurality of instantaneous loudness values. The audio content is filtered by the equalization filter to produce EQ filtered audio content. The plurality of DRC gain values received in the metadata are processed to compute a plurality of DRC gain adjustment values. The plurality of DRC gain adjustment values are applied to the EQ filtered audio content to perform dynamic range control. In another embodiment of dynamic equalization, the computed plurality of instantaneous loudness values are corrected according to the plurality of DRC gain adjustment values to produce corrected instantaneous loudness values, and where the plurality of dynamic parameters defining the equalization filter are computed based on the plurality of corrected instantaneous loudness values. Further, correcting the computed plurality of instantaneous loudness values can include summing the computed plurality of instantaneous loudness values with the plurality of DRC gain adjustment values in dB format. In another aspect, the metadata includes static filter configuration data that specifies one or more of the following for defining the equalization filter: a) type, such as low frequency cut or boost or high frequency cut or boost, b) filter strength, and c) adjustment or effective frequency range. In that case, the equalization filter configured according to the static filter configuration data is dynamically modified by the dynamic parameters while the audio content is passing therethrough. In yet another aspect, computing the plurality of dynamic parameters defining the equalization filter does not use a mix level or a playback level.

[0075] An article of manufacture includes a machine readable medium having stored therein instructions which, when executed by a processor of an audio playback system, perform dynamic audio equalization while applying dynamic range control as follows. Audio content is received, and metadata for the audio content is also received, where the metadata includes a plurality of dynamic range control, DRC, gain values that have been computed for the audio content. The plurality of DRC gain values received in the metadata are applied to inverse DRC characteristics to compute a plurality of instantaneous loudness values for the audio content. A plurality of dynamic parameters defining an equalization filter are computed, where the dynamic parameters are computed based on the computed plurality of instantaneous loudness values. The audio content is filtered by the equalization filter to produce EQ filtered audio content. The plurality of DRC gain values received in the metadata are processed to compute a plurality of DRC gain adjustment values. The plurality of DRC gain adjustment values are applied to the EQ filtered audio content to perform dynamic range control.

[0076] While certain embodiments have been described and shown in the drawings, it is understood that these embodiments are merely for illustrative purposes and not intended to limit the broad disclosure, and that other embodiments will occur to those of ordinary skill in the art to which the present disclosure pertains. Therefore, the descriptions are to be regarded as illustrative in nature and not restrictive.

Claims

1. A method for performing audio equalization in a playback system applying dynamic range control, comprising: a) Receive audio content and metadata for the audio content, wherein the metadata includes multiple dynamic range control (DRC) gain values ​​that have been calculated for the audio content; b) Calculate multiple parameters that define an equalization filter, wherein the parameters are calculated based on the multiple DRC gain values ​​received in the metadata; c) The audio content is filtered by the equalization filter; d) Adjusting the dynamic range of the audio content by applying the plurality of DRC gain values ​​to the audio content, wherein the equalization filter performs filtering compensation on the audio content to adjust the coloring effect of the dynamic range of the audio content; and e) Provide audio content that has been filtered in operation c) and has had its dynamic range adjusted in operation d) to drive the speakers in the playback system.

2. The method according to claim 1, further comprising: Processing the plurality of DRC gain values ​​received in the metadata to calculate a plurality of DRC gain adjustment values, wherein applying the plurality of DRC gain values ​​to the audio content includes applying the DRC gain adjustment values ​​to the audio content.

3. The method of claim 1, wherein the metadata includes static filter configuration data, the static filter configuration data specifying one or more of the following for defining the equalization filter: a) type, including low-frequency cutoff or low-frequency enhancement, or high-frequency cutoff or high-frequency enhancement; b) filter strength; and c) adjusted or effective frequency range. in, The equalizer, configured according to the static filter configuration data, is dynamically modified by the parameters as the audio content passes through it.

4. The method of claim 3, further comprising: Processing the plurality of DRC gain values ​​received in the metadata to calculate a plurality of DRC gain adjustment values, wherein applying the plurality of DRC gain values ​​to the audio content includes applying the DRC gain adjustment values ​​to the audio content.

5. The method of claim 1, wherein the calculation of the plurality of parameters defining the equalization filter does not use a mixing level or a playback level.

6. The method according to claim 1, wherein filtering the audio content by the equalization filter comprises: Bandpass filtering is applied to the background noise in the audio content to prevent the background noise from being amplified.

7. The method of claim 6, wherein applying the plurality of DRC gain values ​​to the audio content comprises: A single-band DRC is applied, wherein the same gain is applied to all frequency components of the audio content.

8. An article comprising: A non-transitory machine-readable medium storing instructions that are executed by a processor of an audio playback system to perform the following operations: a) Receive audio content and metadata for the audio content, wherein the metadata includes multiple dynamic range control (DRC) gain values ​​that have been calculated for the audio content; b) Calculate a plurality of parameters that define an equalization filter, wherein the audio content is filtered by the equalization filter before driving the speakers in the playback system, wherein the parameters are calculated based on the plurality of DRC gain values ​​received in the metadata; c) The audio content is filtered by the equalization filter; d) Adjusting the dynamic range of the audio content by applying the plurality of DRC gain values ​​to the audio content, wherein the equalization filter performs filtering compensation on the audio content to adjust the coloring effect of the dynamic range of the audio content; and e) Provide audio content that has been filtered in operation c) and adjusted in dynamic range in operation d) to drive the speaker in the playback system.

9. The article of manufacture of claim 8, wherein the non-transitory machine-readable medium stores instructions executed by the processor to perform the following operations: processing the plurality of DRC gain values ​​received in the metadata to calculate a plurality of DRC gain adjustment values, wherein applying the plurality of DRC gain values ​​to the audio content includes applying the DRC gain adjustment values ​​to the audio content.

10. The article of manufacture of claim 9, wherein the metadata includes static filter configuration data, the static filter configuration data specifying one or more of the following for defining the equalization filter: a) type, including low-frequency cutoff or low-frequency enhancement, or high-frequency cutoff or high-frequency enhancement; b) filter strength; and c) adjusted or effective frequency range. in, The equalizer, configured according to the static filter configuration data, is dynamically modified by the parameters as the audio content passes through it.

11. The article of manufacture of claim 8, wherein the metadata includes static filter configuration data, the static filter configuration data specifying one or more of the following for defining the equalization filter: a) type, including low-frequency cutoff or low-frequency enhancement, or high-frequency cutoff or high-frequency enhancement; b) filter strength; and c) adjusted or effective frequency range. in, The equalizer, configured according to the static filter configuration data, is dynamically modified by the parameters as the audio content passes through it.

12. The article of manufacture of claim 8, wherein the processor calculates the plurality of parameters defining the equalization filter without using a mixing level or a playback level.

13. The article of manufacture of claim 8, wherein the non-transitory machine-readable medium stores instructions executed by the processor to perform the following operations: configuring the equalization filter as a bandpass filter, the bandpass filter filtering the audio content to avoid amplification of background noise in the audio content.

14. The article of manufacture of claim 8, wherein the non-transitory machine-readable medium stores instructions executed by the processor to perform the following operations: applying the plurality of DRC gain values ​​to the audio content by applying a single-band DRC, wherein the single-band DRC applies the same gain to all frequency components of the audio content.

Citation Information

Patent Citations

  • Metadata for loudness and dynamic range control

    US20140294200A1

  • Loudness adjustment for downmixed audio content

    WO2015038522A1