Method and apparatus for dynamic metadata processing of audio data - Patents.com

JP2024531963A5Pending Publication Date: 2025-09-11DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024509353
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-01
Filing Date
2022-08-24
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing audio playback systems struggle with dynamic loudness adjustment and range compression, particularly in real-time encoding scenarios where the entire audio program's loudness and dynamic range cannot be measured beforehand, leading to altered creative intent and suboptimal playback experiences across different devices.

Method used

A method and apparatus for metadata-based dynamic processing of audio data, involving a decoder that receives a bitstream with audio and metadata, determines processing parameters based on playback conditions, and applies them to adjust loudness and range compression dynamically, using metadata to account for device and environmental constraints.

Benefits of technology

Enables real-time, device-specific loudness management and dynamic range compression, preserving the creative intent of audio content while ensuring consistent playback across various environments and devices, without altering the original audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present application describes a method for metadata-based dynamic processing of audio data for playback, the method comprising: receiving, by a decoder, a bitstream including audio data and metadata for dynamic loudness adjustment; decoding the audio data and metadata to obtain decoded audio data and metadata; determining from the metadata one or more processing parameters for dynamic loudness adjustment based on playback conditions; applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data; and outputting the processed audio data for playback. Further, a method is described for encoding the audio data and metadata for dynamic loudness adjustment into a bitstream. Further, respective data and encoders, respective systems and computer program products are described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates generally to methods of metadata-based dynamic processing of audio data for playback, and in particular to determining and applying one or more processing parameters to audio data for dynamic loudness adjustment and / or dynamic range compression. The present disclosure further relates to methods of encoding audio data and metadata for dynamic loudness adjustment and / or dynamic range compression into a bitstream. The present disclosure also further relates to respective data and encoders, and respective systems and computer program products.

[0002] While certain embodiments are described herein with specific reference to that disclosure, it will be understood that the disclosure is not limited to such fields of use but is applicable in a broader sense. [Background technology]

[0003] Any discussion of background technology throughout this disclosure is in no way to be taken as an admission that such technology is widely known or forms part of the common general knowledge in the field.

[0004] In the reproduction of audio content, loudness is the personal experience of sound pressure. In film or television content, it has been found that the loudness of the dialogue within a program is the most critical parameter determining the listener's perception of program loudness.

[0005] A full program analysis needs to be performed to determine the average loudness of the program, either the whole program or just the dialogue. The average loudness is usually required for loudness compliance (e.g. CALM Act in the US) and is also used to adjust the dynamic range control (DRC) parameters. The dynamic range of a program is the difference between its quietest sound and its loudest sound. The dynamic range of a program depends on its content, e.g. an action movie may have a different, wider dynamic range than a documentary, reflecting the creator's intention. However, the capabilities of devices to reproduce audio content in its original dynamic range vary greatly. In addition to loudness measurement, dynamic range control is thus a further important factor in providing an optimal listening experience.

[0006] To perform loudness management and dynamic range control, the entire audio program or audio program segments should be analyzed and the resulting loudness and DRC parameters may be provided together with the audio data or encoded audio data to be applied at a decoder or playback device.

[0007] When analysis of the entire audio program or audio program segments prior to encoding is not available, e.g., in real-time (dynamic) encoding, loudness processing or leveling is used to ensure loudness compliance, and potential dynamic range limitations are applied, if applicable, depending on the playback requirements. This approach delivers processed audio that is "optimal" for a single playback environment.

[0008] Thus, a need exists for a metadata process that provides the "original" unprocessed audio along with accompanying metadata so that a playback device can use the metadata to dynamically modify the audio in response to device constraints or user requests. Summary of the Invention

[0009] According to a first aspect of the present disclosure, there is provided a method of metadata-based dynamic processing of audio data for playback. The method may include receiving, by a decoder, a bitstream including audio data and metadata for dynamic loudness adjustment. The method may further include decoding, by the decoder, the audio data and the metadata to obtain decoded audio data and the metadata. The method may further include determining, by the decoder, from the metadata, one or more processing parameters for dynamic loudness adjustment based on a playback condition. The method may further include applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data. The method may also include outputting the processed audio data for playback.

[0010] The metadata for dynamic loudness adjustment may include multiple sets of metadata, each set corresponding to a respective (e.g., different) playback condition. In that case, determining the one or more processing parameters for dynamic loudness adjustment from the metadata based on the (particular) playback condition may include selecting a set of metadata corresponding to the (particular) playback condition in response to playback condition information provided to the decoder, and retrieving the one or more processing parameters for dynamic loudness adjustment from the selected set of metadata. In that, the playback condition information may indicate the (particular) playback condition or information derived therefrom.

[0011] In some embodiments, the metadata may indicate processing parameters for dynamic loudness adjustment for multiple playback conditions.

[0012] In some embodiments, determining the one or more processing parameters may further include determining one or more processing parameters for Dynamic Range Compression (DRC) based on the playback conditions.

[0013] In some embodiments, the playback condition information may indicate a particular loudspeaker setup. In general, the playback conditions may include one or more of the following: a device type of the decoder, characteristics of the playback device, characteristics of the loudspeakers, a loudspeaker setup, characteristics of the background noise, characteristics of the ambient noise, and characteristics of the acoustic environment.

[0014] In some embodiments, the selected set of metadata may include a set of DRC sequences (DRCSet). Further, each of the multiple sets of metadata may include a respective set of DRC sequences (DRCSet). In general, determining the one or more processing parameters may be said to further include selecting, by the decoder, at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix that corresponds to the playback conditions.

[0015] In some embodiments, determining the one or more processing parameters may further include identifying a metadata identifier indicative of at least one of the selected DRCSet, EQSet, and downmix to determine the one or more processing parameters from the metadata. Specifically, selecting a set of metadata may include identifying a set of metadata corresponding to a particular downmix. The particular downmix may be determined based on a loudspeaker setup.

[0016] In some embodiments, the metadata may include one or more processing parameters related to the average loudness value and, optionally, one or more processing parameters related to the dynamic range compression characteristics. Specifically, each set of metadata may include one or more such processing parameters related to the average loudness value and, optionally, one or more processing parameters related to the dynamic range compression characteristics.

[0017] In some embodiments, the bitstream may further include additional metadata for static loudness adjustments applied to the decoded audio data.

[0018] In some embodiments, the bitstream may be an MPEG-D DRC bitstream, and the presence of metadata may be signaled based on the MPEG-D DRC bitstream syntax.

[0019] In some embodiments, a loudnessInfoSetExtension() element may be used to carry the metadata as a payload.

[0020] In some embodiments, the metadata may include one or more metadata payloads, each of which may include multiple sets of parameters and identifiers, each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for those identifiers in the set.

[0021] In some embodiments, determining the one or more processing parameters may include selecting a set from among a plurality of sets in the payload based on at least one of the DRCSet, EQSet, and downmix selected by the decoder, and the one or more processing parameters determined by the decoder may be one or more processing parameters for an identifier in the selected set.

[0022] According to a second aspect of the present disclosure, there is provided a decoder for metadata-based dynamic processing of audio data for playback, the decoder may include one or more processors and a non-transitory memory and be configured to perform a method including receiving, by the decoder, a bitstream including audio data and metadata for dynamic loudness adjustment, decoding, by the decoder, the audio data and metadata to obtain decoded audio data and metadata, determining, by the decoder, from the metadata one or more processing parameters for dynamic loudness adjustment based on a playback condition, applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data, and outputting the processed audio data for playback.

[0023] The metadata for dynamic loudness adjustment may include multiple sets of metadata, each set corresponding to a respective (e.g., different) playback condition. In that case, determining the one or more processing parameters for dynamic loudness adjustment from the metadata based on the (particular) playback condition may include selecting a set of metadata corresponding to the (particular) playback condition in response to playback condition information provided to the decoder, and retrieving the one or more processing parameters for dynamic loudness adjustment from the selected set of metadata. In that, the playback condition information may indicate the (particular) playback condition or information derived therefrom.

[0024] According to a third aspect of the present disclosure, there is provided a method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream. The method may include inputting original audio data for loudness processing into a loudness leveller to obtain loudness processed audio data as output from the loudness leveller. The method may further include generating metadata for dynamic loudness adjustment based on the loudness processed audio data and the original audio data. The method may also include encoding the original audio data and the metadata into the bitstream.

[0025] In some embodiments, the metadata may include multiple sets of metadata, each set of metadata may correspond to a respective (e.g., different) playback condition.

[0026] In some embodiments, the method may further include generating additional metadata for static loudness adjustment for use by the decoder.

[0027] In some embodiments, generating the metadata may include comparing the loudness processed audio data to the original audio data, and the metadata may be generated based on a result of the comparison.

[0028] In some embodiments, generating the metadata may further include measuring loudness over one or more predefined periods, and the metadata may be generated further based on the measured loudness.

[0029] In some embodiments, the measuring may include measuring an overall loudness of the audio data.

[0030] In some embodiments, measuring may include measuring a loudness of dialogue in the audio data.

[0031] In some embodiments, the bitstream may be an MPEG-D DRC bitstream, and the presence of metadata may be signaled based on MPEG-D DRC bitstream syntax.

[0032] In some embodiments, the loudnessInfoSetExtension() element may be used to carry the metadata as a payload.

[0033] In some embodiments, the metadata may include one or more metadata payloads, each of which may include multiple sets of parameters and identifiers, each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for those identifiers in the set, where the one or more processing parameters may be parameters for dynamic loudness adjustment by a decoder.

[0034] In some embodiments, at least one of drcSetId, eqSetId, and downmixId may relate to at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix selected by the decoder.

[0035] According to a fourth aspect of the present disclosure, there is provided an encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, the encoder may include one or more processors and a non-transitory memory, and may be configured to perform a method including: inputting original audio data to a loudness leveller for loudness processing to obtain loudness-processed audio data as output from the loudness leveller, generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data, and encoding the original audio data and the metadata into a bitstream.

[0036] According to a fifth aspect of the present disclosure, there is provided a system comprising an encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, and a decoder for metadata dynamic processing of the audio data for playback.

[0037] According to a sixth aspect of the present disclosure, there is provided a computer program product having a computer-readable storage medium comprising instructions adapted, when executed by a device having processing capability, to cause the device to perform a method of metadata dynamic processing of audio data for playback or a method of encoding audio data and metadata for dynamic loudness adjustment into a bitstream.

[0038] According to a seventh aspect of the present disclosure, there is provided a computer readable storage medium storing a computer program product as described herein.

[0039] Exemplary embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]

[0040] [Figure 1]1 illustrates an example decoder for metadata based dynamic processing of audio data for playback. [Diagram 2] 1 illustrates an example method for metadata based dynamic processing of audio data for playback. [Diagram 3] 1 illustrates an example of an encoder that encodes original audio data and metadata for dynamic loudness adjustment into a bitstream. [Figure 4] 1 illustrates an example of a method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream. [Diagram 5] 1 represents an example of a device having one or more processors and non-transitory memory and configured to perform the methods described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0041] [overview] The average loudness of a program or dialogue is the main parameter or value used for loudness compliance of broadcast or streaming programs. The average loudness is usually set to -24 or -23 LKFS. With audio codecs that support loudness metadata, this single loudness value, which represents the loudness of the entire program, is carried in the bitstream. Using this value in the decoding process, gain adjustments can be made that result in a predictable playback level, so that the program is played back at a known and consistent level. It is therefore important that this loudness value is set appropriately and accurately. This is not possible, however, in the case of real-time situations such as dynamic encoding with unknown loudness and dynamic range variations, since the average loudness relies on measurements of the entire program before encoding.

[0042] When it is not possible to measure the loudness of the entire file before encoding, a dynamic loudness leveller is often used to modify or contour the audio data before encoding so that it meets the required loudness. Such loudness management is often considered an inferior method to meet compliance because it often changes the audio content dynamic range correlation and thus potentially alters the creative intent. This is especially true when it is desirable to distribute one audio asset to all playback devices, which is one of the advantages of metadata-driven codecs and distribution systems.

[0043] In some approaches, the audio content is mixed to the required target loudness and the corresponding loudness metadata is set to that value. The loudness leveller may still be used in these situations to help guide the audio content to the target loudness; it is less "aggressive" and is used only when the audio content starts to deviate from the required target loudness.

[0044] In view of the above, the method and device described herein aims to make the real-time processing situation, also called dynamic processing situation, metadata-driven as well, which allows dynamic loudness adjustment and dynamic range compression in the real-time situation. The described method and device advantageously: ● Real-time loudness adjustment and use of DRC in the MPEG-D DRC syntax; Use of real-time loudness adjustment and DRC in combination with downmixId; ● Real-time loudness adjustment and use of DRC in combination with drcSetId; Use of real-time loudness adjustment and DRC in combination with eqSetId This makes it possible.

[0045] That is, depending on the decoder settings (e.g., DRCSet, EQSet, and downmix), the decoder can search a given payload for an appropriate parameter-identifier set based on the syntax by matching said settings with the identifiers. The parameters included in the set whose identifier best matches the settings can then be selected as the processing parameters for the dynamic loudness adjustment to be applied to the received original audio data for correction.

[0046] Additionally, multiple sets of parameters for dynamic processing (multiple instances of dynLoudCompValue) can be transmitted.

[0047] In addition to correcting the overall loudness, metadata-driven dynamic loudness compensation can also be used to "center" the calculation and application of DRC gains. This centering can be a consequence of the correction of the loudness of the content by dynamic loudness compensation and the way DRC is normally calculated and applied. In this sense, it can be said that metadata for dynamic loudness compensation is used to adjust the DRC parameters.

[0048] [Metadata dynamic processing of audio data] With reference to the example of Fig. 1, a decoder 100 for metadata-based dynamic processing of audio data for playback is described. The decoder 100 may have one or more processors and non-transitory memory and may be configured to perform a method including the process represented in the example of Fig. 2 by steps S101-S105.

[0049] The decoder 100 may receive a bitstream containing audio data and metadata and, depending on the requirements, may be able to output unprocessed (original) audio data, processed audio data after application of dynamic processing parameters determined from the metadata, and / or the metadata itself.

[0050] Referring to the example of Fig. 2, in step S101, the decoder 100 may receive a bitstream including audio data and metadata for dynamic loudness adjustment, and optionally dynamic range compression (DRC). The audio data may be encoded audio data, or the audio data may further be unprocessed. That is, the audio data may be said to be original audio data. The metadata may include multiple sets of parameters. For example, each metadata payload may include multiple such sets of metadata. These different sets of metadata may relate to respective playback conditions (e.g., different playback conditions).

[0051] The format of the bitstream is not limited, but in an embodiment, the bitstream may be an MPEG-D DRC bitstream. The presence of metadata for dynamic processing of audio data may then be signaled based on the MPEG-D DRC bitstream syntax. In an embodiment, a loudnessInfoSetExtension() element may be used to carry the metadata as payload, as described further below.

[0052] In step S102, the audio data and metadata may then be decoded by a decoder to obtain decoded audio data and metadata. In an embodiment, the metadata may include one or more processing parameters related to the average loudness value, and optionally one or more processing parameters related to the dynamic range compression characteristics. It is understood that each set of metadata may include a respective processing parameter.

[0053] The metadata allows for the application of dynamic or real-time corrections. For example, in the case of encoding and decoding for live real-time playback, application of "real-time" or dynamic loudness metadata is desired to ensure that the live playback audio is properly loudness managed.

[0054] In step S103, the decoder then determines from the metadata one or more processing parameters for dynamic loudness adjustment based on the playback conditions, which may be done by using the playback conditions or information derived from the playback conditions (e.g., playback condition information) to identify an appropriate set of metadata from among multiple sets of metadata.

[0055] In an embodiment, the playback conditions may include one or more of the following: a decoder device type, playback device characteristics, loudspeaker characteristics, loudspeaker setup, background noise characteristics, ambient noise characteristics, and acoustic environment characteristics. Preferably, the playback condition information may indicate a specific loudspeaker setup. Considering the playback conditions allows the decoder to target the selection of processing parameters for dynamic loudness adjustment according to device and environmental constraints.

[0056] In an embodiment, the process of determining one or more processing parameters in step S130 may further include selecting, by the decoder, at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix corresponding to the playback conditions, whereby at least one of the DRCSet, EQSet, and downmix correlates with or indicates individual device and environmental constraints due to the playback conditions.

[0057] Preferably, step S103 includes selecting a set of DRC sequences (DRCSet), in other words, the selected metadata set may include such a set of DRC sequences.

[0058] In an embodiment, the process of determining in step S103 may further comprise identifying a metadata identifier indicating at least one selected DRCSet, EQSet and DownmixSet in order to determine one or more processing parameters from the metadata. The metadata identifier thus allows associating the metadata with the corresponding selected DRCSet, EQSet and / or downmix and thus with the respective playback conditions.

[0059] In an embodiment, a particular loudspeaker setup may be used to determine a downmix, which may then be used to identify and select an appropriate one among the sets of metadata. In such cases, the particular loudspeaker setup and / or downmix may be indicated by the playback condition information described above.

[0060] In an embodiment, the metadata may include one or more metadata payloads (e.g., dynLoudComp() payloads as shown in Table 5 below), each of which may include multiple sets of parameters (e.g., parameter dynLoudCompValue) and identifiers, with each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for the identifiers in the set. That is, each payload may have an array of entries, with each entry including processing parameters and identifiers (e.g., drcSetId, eqSetId, downmixId). The array of entries may correspond to the multiple sets of metadata described above. Preferably, each entry may include a downmix identifier.

[0061] In a further embodiment, the determining in step S103 may thus comprise selecting a set from among the sets in the payload based on the downmix selected by the decoder (or alternatively based on at least one of the DRCSet, EQSet and downmix), and the one or more processing parameters determined in step S103 may be one or more processing parameters related to an identifier in the selected set. That is, depending on the settings (e.g. DRCSet, EQSet and downmix) present in the decoder, the decoder may search a given payload for a suitable parameter-identifier set by matching said settings with the identifiers. The parameters included in the set whose identifier best matches the settings may then be selected as the processing parameters for the dynamic loudness adjustment to be applied to the received original audio data for correction.

[0062] In step S104, the determined one or more processing parameters may then be applied by the decoder to the decoded audio data to obtain processed audio data, which in this way is properly loudness managed, e.g. live real-time audio data.

[0063] In step S105, the processed audio data may then be output for playback.

[0064] In an embodiment, the bitstream may further include additional metadata for static loudness adjustment applied to the decoded audio data. Static loudness adjustment refers to processing performed for general loudness normalization, as opposed to dynamic processing for real-time situations.

[0065] By carrying metadata for dynamic processing separately from the additional metadata for general loudness normalization, no "real-time" corrections have to be applied.

[0066] For example, in the case of encoding and decoding for live real-time playback, the application of dynamic processing is desired to ensure that the live playback audio is properly loudness managed, but in the case of non-real-time playback, or transcoding where dynamic compensation is not desired or necessary, dynamic processing parameters determined from metadata do not need to be applied.

[0067] By keeping the (dynamic / real-time) metadata for dynamic processing separate from the additional metadata, the original unprocessed content can be preserved if desired. The original audio is encoded along with the metadata. This allows playback devices to selectively apply dynamic processing, and also allows the original audio content to be played on high-end devices that are capable of playing the original audio.

[0068] As mentioned above, there are several advantages to keeping the dynamic loudness metadata separate from long-term loudness measurements / information such as contentLoudness (in ISO / IEC 23003-4): When combined, the content loudness (or what it should be after the dynamic loudness metadata is applied) is not indicative of the actual loudness of the content, as the available metadata becomes a composite value. In addition to removing this ambiguity of what the content loudness (or program or anchor loudness) is, there are several cases where this is particularly beneficial:

[0069] Separating the metadata for dynamic processing allows a decoder or playback device to turn off the application of dynamic processing and instead apply an implemented real-time loudness leveler, avoiding cascading leveling. This situation can arise, for example, if the device's own real-time leveling solution is better than the one used by the audio codec, or, for example, if the device's own real-time leveling solution is always active because it cannot be disabled, causing further processing to reduce resolution and thus impair the playback experience.

[0070] Separating the metadata for dynamic processing further enables transcoding to codecs that do not support dynamic loudness processing and allows custom loudness processing to be applied before re-encoding.

[0071] A further example is a live broadcast with a single encoding for the live feed. Dynamic processing metadata may be used or stored for archive or on-demand services, so that in the latter case more accurate or compliant loudness measurements can be performed on an entire program basis and the appropriate metadata can be reset.

[0072] This can also be beneficial if the use case is one where a fixed target loudness is used throughout the workflow, for example in an R128 compliance situation where -23LKFS is recommended. In this scenario, the addition of dynamic processing metadata is a "safety" measure, the content is expected and close to the desired target, and the addition of dynamic processing metadata is a secondary check. Thus, it is desirable to have the ability to turn it off. The content is expected and close to the desired target, and the addition of dynamic processing metadata is a secondary check. Thus, it is desirable to have the ability to turn it off.

[0073] [Encoding audio data and metadata for dynamic loudness adjustment] With reference to the examples of Figures 3 and 4, an encoder for encoding original audio data and metadata for dynamic loudness adjustment, and optionally dynamic range compression (DRC), into a bitstream is described, the encoder having one or more processors and non-transient memory and may be configured to perform a method including the process represented in the example steps of Figure 4.

[0074] In step S201, original audio data may be input to the loudness leveller 201 for loudness processing, so as to obtain loudness-processed audio data as output from the loudness leveller 201.

[0075] In step S202, metadata for dynamic loudness adjustment may then be generated based on the loudness processed audio data and the original audio data. Appropriate smoothing and time frames may be used to reduce artifacts.

[0076] In an embodiment, step S202 may include comparing the loudness processed audio data with the original audio data by the analyzer 202. The metadata thus generated can emulate the effect of a leveller at the decoder site. The metadata includes: Gain (wideband and / or multiband) processing parameters that, when applied to the original audio, produce loudness-compliant audio for playback; ● Processing parameters that indicate the dynamics of the audio, e.g. Peak-sample and true peak ○Short-term loudness value ○Short-term loudness value changes may include:

[0077] In an embodiment, step S202 may further include measuring, by the analyzer 202, the loudness over one or more predefined time periods, and the metadata may be generated further based on the measured loudness. In an embodiment, the measuring may include measuring an overall loudness of the audio data. Alternatively, or additionally, in an embodiment, the measuring may include measuring the loudness of dialogue in the audio data.

[0078] In step S203, the original audio data and the metadata may then be encoded into a bitstream. The format of the bitstream is not limited, but in an embodiment, the bitstream may be an MPEG-D DRC bitstream, and the presence of the metadata may then be signaled based on the MPEG-D DRC bitstream syntax. In this case, in an embodiment, a loudnessInfoSetExtension() element may be used to carry the metadata as payload, as described in further detail below.

[0079] In an embodiment, the metadata may include one or more metadata payloads, each of which may include multiple sets of parameters and identifiers, each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for the identifiers in the set, which may be parameters for dynamic loudness adjustment by the decoder. In this case, in an embodiment, at least one of drcSetId, eqSetId, and downmixId may relate to at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix selected by the decoder. In general, the metadata may be said to include multiple sets of metadata, each set corresponding to a respective playback condition (e.g., a different playback condition).

[0080] In an embodiment, the method may further comprise generating additional metadata for static loudness adjustment for use by the decoder. Separating the metadata for dynamic loudness processing and the additional metadata in the bitstream, and further encoding the original audio data in the bitstream, has several advantages, as detailed above.

[0081] The methods described herein may be implemented by a decoder or an encoder, respectively, which may comprise one or more processors and non-transitory memory and be configured to perform the above methods. An example of a device with such processing capabilities is depicted in the example of Figure 5, which shows such a device 300 including two processors 301 and a non-transitory memory 302.

[0082] It should be noted that the methods described herein may further be performed in a system having an encoder for encoding original audio data and metadata for dynamic loudness adjustment, and optionally dynamic range compression (DRC), into a bitstream, as described herein, and a decoder for metadata dynamic processing of the audio data for playback.

[0083] The method may further be implemented as a computer program product having a computer readable storage medium comprising instructions adapted, when executed by a device having processing capabilities, to cause the device to perform the method. The computer program product may be stored in the computer readable storage medium.

[0084] [MPEG-D DRC Modified Bitstream Syntax] In the following it is described how the MPEG-D DRC bitstream syntax as described in ISO / IEC 23003-4 can be modified according to the embodiments described herein.

[0085] The MPEG-D DRC syntax can be extended, such as with the loudnessInfoSetExtension() element shown in Table 2 below, to also carry dynamic processing metadata as frame-based dynLoudComp updates.

[0086] For example, another switch case UNIDRCLOUDEXT_DYNLOUDCOMP may be added to the loudnessInfoSetExtension() element as shown in Table 1. The switch case UNIDRCLOUDEXT_DYNLOUDCOMP may be used to specify a new element dynLoudComp() as shown in Table 5. The loudnessInfoSetExtension() element may be an extension of the loudnessInfoSet() element as shown in Table 2. Furthermore, the loudnessInfoSet() element may be part of the uniDRC() element as shown in Table 3. [Table 1] [Table 2] [Table 3] [Table 4] New dynLoudComp(): [Table 5] ● drcSetId allows dynLoudComp (metadata-wise) to be applied per DRC set. - eqSetId allows dynLoudComp to be applied in combination with different settings of the equalization tool. downmixId allows dynLoudComp to be applied per DownmixId.

[0087] In some cases, in addition to the above parameters, it may be beneficial for the dynLoudComp() element to also include a methodDefinition parameter (e.g., specified by 4 bits) that specifies the loudness measurement method used to derive the dynamic program loudness metadata (e.g., anchor loudness, program loudness, short-term parameters, momentary loudness, etc.) and / or a measurementSystem parameter (e.g., specified by 4 bits) that specifies the loudness measurement system used to measure the dynamic program loudness metadata (e.g., EBU R.128, ITU-R BS-1770 with or without preprocessing, ITU-R BS-1771, etc.). Such parameters may, for example, be included in the dynLoudComp() element between the downmixId and dynLoudCompValue parameters.

[0088] [Alternative Syntax 1] [Table 6] [Table 7] [Table 8] In some cases, it may be beneficial to modify the syntax shown above in Table 8 so that the dynLoudCompPresent parameter and (if dynLoudCompPresent==1) the dynLoudCompValue parameter follow the reliability parameter within the measurementCount loop of loudnessInfoV2(), rather than being outside the measurementCount loop. Additionally, it may also be beneficial to set dynLoudCompValue equal to 0 when dynLoudCompPresent is 0.

[0089] [Alternative Syntax 2] Alternatively, the dynLoudComp() element may be placed in the uniDrcGainExtension(). [Table 9] [Table 10] [Table 11] Semantics dynLoudCompValue: This field contains the value of dynLoudCompDb. The value is encoded according to the following table. The default value is 0 dB. [Table 12] [Updated MPEG-D DRC loudness normalization processing] [Table 13] [Pseudocode for dynLoudComp selection and processing]

number

number

number

number

[0090] [Alternative updated MPEG-D DRC loudness normalization process] [Table 14] If the alternative loudness normalization process of Table 14 above is used, the loudness normalization process pseudocode described above may be replaced by the following alternative loudness normalization process pseudocode: Note that a default value of dynLoudCompDb, e.g., 0 dB, may be assumed to ensure that the value of dynLoudCompDb is defined even if no dynamic loudness processing metadata is present in the bitstream.

number

[0091] targetLoudness: This field contains the desired output loudness. The value is coded according to the following table: [Table 22] dynLoudnessNormalizationOn: This flag tells whether the dynamic loudness normalization process is turned on or off. The default value is 0. If dynLoudnessNormalizationOn==0, dynloudnessNormalizationGainDb should be set to 0dB.

[0092] [interpretation] Unless otherwise stated, and as will be apparent from the discussion that follows, throughout this disclosure, discussions using terms such as "processing," "computing," "determining," "analyzing," and the like will be understood to refer to the operation and / or processing of a computer or computing system or similar electronic computing device that manipulates and / or converts data represented as physical quantities, such as electrons, into other data also represented as physical quantities.

[0093] In a similar manner, the term "processor" may refer to any device or part of a device that processes electronic data, e.g., from registers and / or memory, and converts the electronic data into other electronic data, which may be stored, e.g., in registers and / or memory. A "computer" or "computing machine" or "computing platform" may include one or more processors.

[0094] The methodologies described herein, in an example embodiment, are executable by one or more processors that accept computer-readable (also referred to as machine-readable) code that includes a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed is included. Thus, an example is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes a main RAM and / or static RAM, and / or a ROM. A bus subsystem may be included for communication between components. The processing system may also be a distributed processing system that includes processors coupled by a network. If the processing system requires a display, such a display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system also includes an input device, such as one or more of an alphanumeric input unit, such as a keyboard, a pointing and control device, such as a mouse, and the like. The processing system may also include a storage system, such as a disk drive unit. The processing system in some configurations may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that carries computer-readable code (e.g., software) that includes a set of instructions that, when executed by one or more processors, causes the execution of one or more of the methods described herein. It should be noted that, when a method includes several elements, e.g. several steps, no order of such elements is implied unless specifically stated. The software may reside on a hard disk, or may reside, completely or at least partially, in a RAM and / or in a processor during its execution by the computer system.Thus, the memory and the processor also constitute a computer-readable carrier medium carrying computer-readable code, which may further form or be included in a computer program product.

[0095] In alternative, exemplary embodiments, one or more processors may operate as stand-alone devices or may be connected, e.g., networked, to other processors, where in a networked arrangement, one or more processors may operate in the capacity of a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular telephone, web appliance, network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed by the machine.

[0096] Additionally, the term "machine" should also be interpreted as including any collection of machines that execute a set (or sets) of instructions, either individually or collectively, to perform any one or more of the methodologies discussed herein.

[0097] Thus, an exemplary embodiment of each of the methods described herein takes the form of a computer-readable carrier medium carrying a set of instructions, e.g., a computer program executed on one or more processors, e.g., one or more processors that are part of a web server arrangement. Thus, as will be appreciated by one of ordinary skill in the art, the exemplary embodiments of the present disclosure may be embodied as a method, an apparatus, such as a special purpose device, an apparatus, such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code, which when executed on one or more processors, causes the one or more processors to perform the method. Accordingly, aspects of the present disclosure may take the form of a method, an entirely hardware exemplary embodiment, an entirely software exemplary embodiment, or an exemplary embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0098] The software may also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in an exemplary embodiment, the term "carrier medium" should be interpreted as including a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more instructions. The term "carrier medium" should also be interpreted as including any medium capable of storing, encoding, or carrying a set of instructions for execution by one or more of the processors, causing the one or more processors to execute any one or more of the methodologies disclosed herein. The carrier medium may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and optical-magnetic disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wiring, and optical fibers, including the wires that comprise a bus subsystem. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should be understood to include, but is not limited to, computer products embodied in solid-state memories, optical media, and magnetic media, media carrying propagated signals detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, perform a method, and transmission media within a network carrying propagated signals detectable by at least one processor of one or more processors and representing a set of instructions.

[0099] It will be understood that the steps of the methods discussed are, in one example embodiment, performed by a suitable processor(s) of a processing (e.g., computer) system executing instructions (computer-readable code) stored in storage. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.

[0100] References throughout this disclosure to "one embodiment," "some embodiments," or "exemplary embodiments" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this disclosure. Thus, the appearances of the phrases "in one embodiment," "in some embodiments," or "in exemplary embodiments" in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiments. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more exemplary embodiments.

[0101] As used herein, unless otherwise specified, the use of ordinal adjectives such as "first," "second," "third," etc. to describe a common object merely indicates that different instances of the same object are being referred to, and is not intended to imply that the objects so described must be in a given order temporally, spatially, ranked, or otherwise.

[0102] In the following claims and the description herein, any one of the terms "comprising", "comprised of" or "which comprises" is an open term meaning to include at least the following element / feature but not to exclude others. Thus, the term "comprising", when used in the claims, should not be interpreted as being limited to the means or elements or steps listed thereafter. For example, the scope of the expression "a device comprising A and B" should not be limited to a device consisting of only elements A and B. Any one of the terms "including" or "which includes" as used herein is also an open term meaning to include at least the following element / feature but not to exclude others. Thus, "comprising" is synonymous with and means "having".

[0103] In the above description of exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure are sometimes grouped together in a single exemplary embodiment, figure, or description thereof in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. It should be noted, however, that this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single above-disclosed exemplary embodiment. Thus, the claims following the specification are hereby incorporated into this specification, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.

[0104] Moreover, while some exemplary embodiments described herein include some features included in other exemplary embodiments but not others, combinations of features of different exemplary embodiments are intended to be within the scope of the present disclosure and form different exemplary embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed exemplary embodiments can be used in any combination.

[0105] In the description provided herein, numerous specific details are set forth. However, it will be understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure an understanding of this specification.

[0106] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other changes and further modifications may be made thereto without departing from the spirit of the disclosure, and it is intended to claim all such changes and modifications as being within the scope of the disclosure. For example, any formulas given above are merely representative of procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the disclosure.

[0107] The following enumerated example embodiments (EEE) describe certain structures, features, and functions of certain aspects of example implementations disclosed herein.

[0108] EEE1. 1. A method for metadata based dynamic processing of audio data for playback, comprising: (a) receiving, by a decoder, a bitstream including audio data and metadata for dynamic loudness adjustment; (b) decoding, by the decoder, the audio data and the metadata to obtain decoded audio data and the metadata; and (c) determining, by the decoder, one or more processing parameters for dynamic loudness adjustment based on the playback condition information; and (d) applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data; and (e) outputting the processed audio data for playback; The method according to claim 1,

[0109] EEE2. the metadata indicating processing parameters for dynamic loudness adjustment for a plurality of playback conditions; The method described in EEE1.

[0110] EEE3. determining the one or more processing parameters further comprises determining one or more processing parameters for dynamic range compression (DRC) based on the playback conditions. The method according to EEE1 or EEE2.

[0111] EEE4. the playback conditions include one or more of the following: a device type of the decoder, a playback device characteristic, a loudspeaker characteristic, a loudspeaker setup, a background noise characteristic, an ambient noise characteristic, and an acoustic environment characteristic; A method according to any one of EEE1 to EEE3.

[0112] EEE5. The process (c) further includes selecting, by the decoder, any one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix corresponding to the playback condition; A method according to any one of EEE1 to EEE4.

[0113] EEE6. The process (c) further includes identifying a metadata identifier indicating the at least one selected DRCSet, EQSet, and downmix to determine the one or more processing parameters from the metadata. The method described in EEE5.

[0114] EEE7. the metadata includes one or more processing parameters related to average loudness values ​​and, optionally, one or more processing parameters related to dynamic range compression characteristics; A method according to any one of EEE1 to EEE6.

[0115] EEE8. the bitstream further comprises additional metadata for static loudness adjustment applied to the decoded audio data. A method according to any one of EEE1 to EEE7.

[0116] EEE9. the bitstream is an MPEG-D DRC bitstream and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; A method according to any one of EEE1 to EEE8.

[0117] EEE10. The loudnessInfoSetExtension() element is used to carry the metadata as payload. The method described in EEE9.

[0118] EEE11. the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for identifiers in the set; 13. The method of any one of EEE1 to EEE10.

[0119] EEE12. The process (c) includes selecting a set from among a plurality of sets in the payload based on at least one of the DRCSet, the EQSet, and the downmix selected by the decoder; the one or more process parameters determined in process (c) are one or more process parameters related to identifiers in the selected set; The method described in EEE11 which is subordinate to EEE5.

[0120] EEE13. 1. A decoder for metadata based dynamic processing of audio data for playback, the decoder comprising one or more processors and a non-transitory memory, (a) receiving, by a decoder, a bitstream including audio data and metadata for dynamic loudness adjustment; (b) decoding, by the decoder, the audio data and the metadata to obtain decoded audio data and the metadata; and (c) determining, by the decoder, one or more processing parameters for dynamic loudness adjustment based on the playback condition information; and (d) applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data; and (e) outputting the processed audio data for playback; 23. A decoder configured to perform a method comprising:

[0121] EEE14. 1. A method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream, comprising the steps of: (a) inputting original audio data to a loudness leveller for loudness processing to obtain loudness-processed audio data as output from the loudness leveller; (b) generating metadata for the dynamic loudness adjustment based on the loudness processed audio data and the original audio data; and (c) encoding the original audio data and the metadata into the bitstream; and The method according to claim 1,

[0122] EEE15. generating additional metadata for static loudness adjustment for use by the decoder. Method according to EEE14.

[0123] EEE16. and process (b) includes comparing the loudness processed audio data with the original audio data, and the metadata is generated based on a result of the comparison. The method according to claim EEE14 or EEE15.

[0124] EEE17. and process (b) further comprises measuring loudness over one or more predefined time periods, and the metadata is generated further based on the measured loudness. Method according to EEE16.

[0125] EEE18. said measuring comprises measuring an overall loudness of said audio data. Method according to EEE17.

[0126] EEE19. said measuring comprising measuring loudness of dialogue in said audio data; Method according to EEE17.

[0127] EEE20. the bitstream is an MPEG-D DRC bitstream and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; 8. The method according to any one of claims 8 to 9.

[0128] EEE21. The loudnessInfoSetExtension() element is used to carry the metadata as payload. The method described in EEE20.

[0129] EEE22. the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including at least one of a DRCSet identifier (drcSetId), an EQSet identifier (eqSetId), and a downmix identifier (downmixId) in combination with one or more processing parameters for identifiers in the set, the one or more processing parameters being parameters for dynamic loudness adjustment by a decoder; 13. The method according to any one of claims 8 to 12.

[0130] EEE23. At least one of drcSetId, eqSetId, and downmixId relates to at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix selected by the decoder; The method described in EEE22.

[0131] EEE24. 1. An encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, the encoder comprising one or more processors and a non-transitory memory, the encoder comprising: (a) inputting original audio data to a loudness leveller for loudness processing to obtain loudness-processed audio data as output from the loudness leveller; (b) generating metadata for the dynamic loudness adjustment based on the loudness processed audio data and the original audio data; and (c) encoding the original audio data and the metadata into the bitstream; and 23. An encoder configured to perform a method comprising:

[0132] EEE25. An encoder for encoding original audio data and metadata for dynamic loudness adjustment and dynamic range compression (DRC) into a bitstream, as claimed in claim 24; A decoder for metadata-based dynamic processing of audio data for playback according to EEE13; A system having

[0133] EEE26. A computer program product having a computer-readable medium comprising instructions adapted, when executed by a device having processing capability, to cause the device to perform a method according to any one of EEE1 to EEE12 or EEE14 to EEE23.

[0134] EEE27. A computer readable storage medium storing a computer program product according to EEE26.

[0135] EEE28. receiving, by the decoder, via an interface, an indication of whether to perform the metadata database dynamic processing of audio data for playback; and bypassing the step of applying at least the determined one or more processing parameters to the decoded audio data if the decoder receives an instruction not to perform the metadata database dynamic processing of the audio data for playback. The method of any one of EEE1 to EEE12, further comprising:

[0136] EEE29. the decoder bypasses at least the step of applying the determined one or more processing parameters to the decoded audio data until the decoder receives, via the interface, the indication of whether to perform the metadata database dynamic processing of the audio data for playback. Method according to EEE28.

[0137] EEE30. the metadata indicates a plurality of processing parameters for dynamic loudness adjustment for a plurality of playback conditions; the metadata further comprises parameters specifying a loudness measurement method used to derive a processing parameter in the plurality of processing parameters. 2. A method according to any one of EEE1 to EEE12, EEE28, or EEE29.

[0138] EEE31. the metadata indicates a plurality of processing parameters for dynamic loudness adjustment for a plurality of playback conditions; the metadata further comprises a parameter specifying a loudness measurement system used to measure a processing parameter in the plurality of processing parameters. 2. A method according to any one of EEE1 to EEE12 or EEE28 to EEE30.

[0139] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to European Patent Application No. 21193209.0, filed August 26, 2021, and U.S. Provisional Patent Application No. 63 / 237231, filed August 26, 2021, and U.S. Provisional Patent Application No. 63 / 251307, filed October 1, 2021, all of which are incorporated by reference in their entirety into this application.

Claims

1. 1. A method for metadata-based dynamic processing of audio data for playback, comprising: receiving, by a decoder, a bitstream including audio data and metadata for dynamic loudness adjustment, the metadata for dynamic loudness adjustment including a plurality of sets of metadata, each set of metadata corresponding to a respective playback condition; decoding, by the decoder, the audio data and the metadata to obtain decoded audio data and the metadata; selecting a set of metadata corresponding to a particular playback condition in response to playback condition information provided to the decoder, and deriving from the selected set of metadata one or more processing parameters for dynamic loudness adjustment; applying the retrieved one or more processing parameters to the decoded audio data to obtain processed audio data; and outputting the processed audio data for playback; and the selected set of metadata includes a set of dynamic range compression (DRC) sequences (DRCSet); the bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including a respective downmix identifier (downmixId) in combination with one or more processing parameters for the downmix identifier in the set; method.

2. The loudnessInfoSetExtension() element is used to carry the metadata as payload. The method of claim 1.

3. 1. A decoder for metadata-based dynamic processing of audio data for playback, the decoder comprising one or more processors and a non-transitory memory, The one or more processors: receiving, by the decoder, a bitstream including audio data and metadata for dynamic loudness adjustment, the metadata for dynamic loudness adjustment including a plurality of sets of metadata, each set of metadata corresponding to a respective playback condition; decoding, by the decoder, the audio data and the metadata to obtain decoded audio data and the metadata; selecting a set of metadata corresponding to a particular playback condition in response to playback condition information provided to the decoder, and deriving from the selected set of metadata one or more processing parameters for dynamic loudness adjustment; applying the retrieved one or more processing parameters to the decoded audio data to obtain processed audio data; and outputting the processed audio data for playback; configured to perform a method comprising: the selected set of metadata includes a set of dynamic range compression (DRC) sequences (DRCSet); the bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including a respective downmix identifier (downmixId) in combination with one or more processing parameters for the downmix identifier in the set; decoder.

4. 1. A method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream, comprising: inputting original audio data to a loudness leveler for loudness processing to obtain loudness-processed audio data as output from the loudness leveler; generating metadata for the dynamic loudness adjustment based on the loudness-processed audio data and the original audio data; encoding the original audio data and the metadata into the bitstream; and The metadata includes a plurality of sets of metadata, each set of metadata corresponding to a respective playback condition; the bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including a respective downmix identifier (downmixId) in combination with one or more processing parameters for the downmix identifier in the set, the one or more processing parameters being parameters for dynamic loudness adjustment by a decoder; method.

5. generating additional metadata for static loudness adjustment for use by the decoder. The method of claim 4.

6. generating the metadata includes comparing the loudness-processed audio data with the original audio data, and the metadata is generated based on a result of the comparison. The method of claim 4.

7. generating the metadata further includes measuring loudness over one or more predefined time periods, and the metadata is generated further based on the measured loudness. The method of claim 6.

8. said measuring comprising measuring an overall loudness of said audio data. The method of claim 7.

9. said measuring comprising measuring the loudness of dialogue in said audio data. The method of claim 8.

10. The loudnessInfoSetExtension() element is used to carry the metadata as payload. The method of claim 4.

11. 1. An encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, the encoder comprising one or more processors and a non-transitory memory, The one or more processors: inputting original audio data to a loudness leveler for loudness processing to obtain loudness-processed audio data as output from the loudness leveler; generating metadata for the dynamic loudness adjustment based on the loudness-processed audio data and the original audio data; encoding the original audio data and the metadata into the bitstream; configured to perform a method comprising: The metadata includes a plurality of sets of metadata, each set of metadata corresponding to a respective playback condition; the bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled according to MPEG-D DRC bitstream syntax; the metadata includes one or more metadata payloads, each metadata payload including a plurality of sets of parameters and identifiers, each set including a respective downmix identifier (downmixId) in combination with one or more processing parameters for the downmix identifier in the set, the one or more processing parameters being parameters for dynamic loudness adjustment by a decoder; Encoder.

12. an encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream according to claim 11; A decoder for metadata-based dynamic processing of audio data for playback according to claim 3, A system having:

13. A computer program comprising instructions adapted, when executed by a device having processing capabilities, to cause said device to carry out the method of any one of claims 1 to 2 or 4 to 10.

14. A computer-readable storage medium storing the computer program of claim 13.