Method and apparatus for dynamic processing of audio data metadata databases
Patent Information
- Application Number
- JP2024509353
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-01
- Filing Date
- 2022-08-24
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2042-08-24
Smart Images

Figure 0007927834000028 
Figure 0007927834000029 
Figure 0007927834000030
Abstract
Description
[[Technical Field]]
[0001] The present disclosure generally relates to a method for dynamic processing of metadata of audio data for playback, and specifically relates to determining one or more processing parameters for dynamic loudness adjustment and / or dynamic range compression and applying the parameters to audio data. The present disclosure further relates to a method for encoding audio data and metadata for dynamic loudness adjustment and / or dynamic range compression into a bitstream. The present disclosure also further relates to respective data and encoders, and respective systems and computer program products.
[0002] While some embodiments are described herein with specific reference to the disclosure, it will be understood that the present disclosure is not limited to such fields of use and is applicable in a broader sense. [[Background Art]]
[0003] Any discussion of the background art throughout the present disclosure shall in no way be regarded as an admission that such technology was widely known or formed part of the common general knowledge in the art.
[0004] In the playback of audio content, loudness is a personal experience of sound pressure. In movie or television content, it has been found that the loudness of dialog (conversation or talk) within a program is the most critical parameter that determines a listener's perception of program loudness.
[0005] To determine the average loudness of a program, either the entire program or only the dialogue, an analysis of the entire program must be performed. Average loudness is typically required for loudness compliance (e.g., the US CALM law) and is also used to adjust Dynamic Range Control (DRC) parameters. The dynamic range of a program is the difference between its quietest and loudest sounds. The dynamic range of a program depends on its content; for example, an action movie may have a wider dynamic range than a documentary, reflecting the creator's intentions. However, the ability of devices to reproduce audio content at its original dynamic range varies greatly. In addition to loudness measurement, dynamic range control is thus another crucial factor in providing an optimal listening experience.
[0006] To perform loudness management and dynamic range control, the entire audio program or a segment of the audio program should be analyzed, and the resulting loudness and DRC parameters can be supplied along with the audio data or encoded audio data to be applied by the decoder or playback device.
[0007] When analysis of the entire audio program or audio program segment before encoding is not available, for example in real-time (dynamic) encoding, loudness processing or leveling is used to ensure loudness compliance, and, where applicable, potential dynamic range limitations are applied depending on playback requirements. This approach provides processed audio that is "optimal" for a single playback environment.
[0008] Therefore, there is a need for metadata processing that provides the “original” raw audio along with the accompanying metadata, so that playback devices can use metadata to dynamically modify the audio in response to device constraints or user requests. [Overview of the Initiative]
[0009] A method for metadata dynamic processing of audio data for playback is provided in accordance with the first aspect of this disclosure. The method may include a decoder receiving a bitstream containing audio data and metadata for dynamic loudness adjustment. The method may further include the decoder decoding the audio data and metadata to obtain decoded audio data and metadata. The method may further include the decoder determining one or more processing parameters for dynamic loudness adjustment from the metadata based on playback conditions. The method may further include applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data. The method may also include outputting the processed audio data for playback.
[0010] The metadata for dynamic loudness adjustment may include multiple sets of metadata, each set corresponding to a different (e.g., different) playback condition. In this case, determining one or more processing parameters for dynamic loudness adjustment from the metadata based on a (specific) playback condition may include selecting a set of metadata corresponding to the (specific) playback condition in response to playback condition information supplied to the decoder, and retrieving one or more processing parameters for dynamic loudness adjustment from the selected set of metadata. In this process, the playback condition information may indicate a (specific) playback condition or information derived therefrom.
[0011] In some embodiments, the metadata may indicate processing parameters for dynamic loudness adjustment for multiple playback conditions.
[0012] In some embodiments, determining one or more processing parameters may further include determining one or more processing parameters for Dynamic Range Compression (DRC) based on playback conditions.
[0013] In some embodiments, playback condition information may indicate a specific loudspeaker setup. Generally, playback conditions may include one or more of the following: decoder device type, playback device characteristics, loudspeaker characteristics, loudspeaker setup, background noise characteristics, ambient noise characteristics, and acoustic environment characteristics.
[0014] In some embodiments, the selected metadata set may include a set of DRC sequences (DRCSet). Furthermore, each of the multiple metadata sets may include a set of DRC sequences (DRCSet). Generally, determining one or more processing parameters may be said to further include the decoder selecting at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix that corresponds to the playback conditions.
[0015] In some embodiments, determining one or more processing parameters may further include identifying metadata identifiers that indicate at least one selected DRCSet, EQSet, and downmix in order to determine one or more processing parameters from the metadata. Specifically, selecting a set of metadata may include identifying a set of metadata that corresponds to a particular downmix. A particular downmix may be determined based on the loudspeaker setup.
[0016] In some embodiments, the metadata may include one or more processing parameters relating to the average loudness value, and optionally, one or more processing parameters relating to the dynamic range compression characteristics. Specifically, each set of metadata may include one or more such processing parameters relating to the average loudness value, and optionally, one or more processing parameters relating to the dynamic range compression characteristics.
[0017] In some embodiments, the bitstream may further include additional metadata for static noise adjustments applied to the decoded audio data.
[0018] In some embodiments, the bitstream may be an MPEG-D DRC bitstream, and the presence of metadata may be signaled based on the MPEG-D DRC bitstream syntax.
[0019] In some embodiments, the loudnessInfoSetExtension() element may be used to carry the metadata as a payload.
[0020] In some embodiments, the metadata may include one or more metadata payloads, each metadata payload may include multiple sets of parameters and identifiers, each set including at least one of the DRCSet identifier (drcSetId), EQSet identifier (eqSetId), and downmix identifier (downmixId) in combination with one or more processing parameters relating to those identifiers within the set.
[0021] In some embodiments, determining one or more processing parameters may include selecting a set from among several sets in the payload based on at least one DRCSet, EQSet, and downmix selected by the decoder, and the one or more processing parameters determined by the decoder may be one or more processing parameters relating to identifiers in the selected set.
[0022] In accordance with a second aspect of this disclosure, a decoder for dynamic processing of a metadata database of audio data for playback is provided. The decoder includes one or more processors and non-temporary memory and may be configured to perform a method including receiving a bitstream containing audio data and metadata for dynamic loudness adjustment; decoding the audio data and metadata to obtain decoded audio data and metadata; determining one or more processing parameters for dynamic loudness adjustment from the metadata based on playback conditions; applying the determined one or more processing parameters to the decoded audio data to obtain processed audio data; and outputting the processed audio data for playback.
[0023] The metadata for dynamic loudness adjustment may include multiple sets of metadata, each set corresponding to a different (e.g., different) playback condition. In this case, determining one or more processing parameters for dynamic loudness adjustment from the metadata based on a (specific) playback condition may include selecting a set of metadata corresponding to the (specific) playback condition in response to playback condition information supplied to the decoder, and retrieving one or more processing parameters for dynamic loudness adjustment from the selected set of metadata. In this process, the playback condition information may indicate a (specific) playback condition or information derived therefrom.
[0024] According to a third aspect of the present disclosure, there is provided a method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream. The method may comprise inputting original audio data to a loudness leveler for loudness processing, so as to obtain loudness-processed audio data as an output from the loudness leveler. The method may further comprise generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data. The method may also comprise encoding the original audio data and the metadata into the bitstream.
[0025] In some embodiments, the metadata may comprise a plurality of sets of metadata. Each set of metadata may correspond to a respective (e.g., different) playback condition.
[0026] In some embodiments, the method may further comprise generating additional metadata for static loudness adjustment to be used by a decoder.
[0027] In some embodiments, generating the metadata may comprise comparing the loudness-processed audio data with the original audio data, and the metadata may be generated based on a result of the comparison.
[0028] In some embodiments, generating the metadata may further comprise measuring loudness over one or more predefined time periods, and the metadata may be further generated based on the measured loudness.
[0029] In some embodiments, the measuring may comprise measuring an overall loudness of the audio data.
[0030] In some embodiments, the measuring may comprise measuring a loudness of dialog in the audio data.
[0031] In some embodiments, the bitstream may be an MPEG-D DRC bitstream, and the presence of metadata may be signaled based on the MPEG-D DRC bitstream syntax.
[0032] In some embodiments, the loudnessInfoSetExtension() element may be used to carry metadata as a payload.
[0033] In some embodiments, the metadata may include one or more metadata payloads, each metadata payload may include multiple sets of parameters and identifiers, each set including at least one of the DRCSet identifier (drcSetId), EQSet identifier (eqSetId), and downmix identifier (downmixId) in combination with one or more processing parameters relating to those identifiers within the set, the one or more processing parameters may be parameters for dynamic loudness adjustment by the decoder.
[0034] In some embodiments, at least one of drcSetId, eqSetId, and downmixId may relate to at least one of a set of DRC sequences selected by the decoder (DRCSet), a set of equalizer parameters (EQSet), and a downmix.
[0035] In accordance with a fourth aspect of this disclosure, an encoder is provided for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream. The encoder may include one or more processors and non-temporary memory and be configured to perform a method including inputting original audio data to a loudness leveler for loudness processing so that loudness-processed audio data is obtained as output from the loudness leveler; generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data; and encoding the original audio data and metadata into a bitstream.
[0036] In accordance with the fifth aspect of this disclosure, a system is provided having an encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, and a decoder for metadata database dynamic processing of the audio data for playback.
[0037] In accordance with the sixth aspect of this disclosure, a computer program product is provided having a computer-readable storage medium, which includes instructions adapted to cause a device having processing capabilities to perform a method of metadata database dynamic processing of audio data for playback, or a method of encoding audio data and metadata for dynamic loudness adjustment into a bitstream.
[0038] A computer-readable storage medium storing a computer program product described herein is provided in accordance with a seventh aspect of this disclosure.
[0039] Hereafter, exemplary embodiments of the present disclosure will be described by reference to the accompanying drawings, merely as examples. [Brief explanation of the drawing]
[0040] [Figure 1]This illustrates an example of a decoder for dynamic processing of a metadata database for audio data playback. [Figure 2] This illustrates an example of a method for dynamically processing a metadata database of audio data for playback. [Figure 3] This shows an example of an encoder that encodes the original audio data and metadata for dynamic loudness adjustment into a bitstream. [Figure 4] This illustrates an example of how to encode audio data and metadata for dynamic loudness adjustment into a bitstream. [Figure 5] This represents an example of a device having one or more processors and non-temporary memory, and configured to perform the methods described herein. [Modes for carrying out the invention]
[0041] [overview] The average loudness of a program or dialogue is a key parameter or value used for loudness compliance in broadcast or streaming programs. The average loudness is typically set to -24 or -23 LKFS. According to audio codecs that support loudness metadata, this single loudness value, representing the overall loudness of the program, is carried in the bitstream. Using this value in the decoding process allows for gain adjustments that result in predictable playback levels, thereby ensuring the program is played back at a known, consistent level. Therefore, it is crucial that this loudness value is set appropriately and accurately. However, since the average loudness depends on measurements of the entire program before encoding, this is not possible in real-time situations such as dynamic encoding with unknown loudness and dynamic range fluctuations.
[0042] When it's not possible to measure the overall loudness of a file before encoding, dynamic loudness levelers are often used to modify or contour audio data before encoding to meet the required loudness. Such loudness management is often considered a poor method for satisfying compliance because it frequently alters the audio content dynamic range correlation, thus potentially altering the creative intent. This is especially true when it's desirable to distribute a single audio asset to all playback devices, which is one of the advantages of metadata-driven codecs and distribution systems.
[0043] In some approaches, audio content is mixed with the required target loudness, and the corresponding loudness metadata is set to that value. Loudness levelers can still be used in these situations as well, as they are used to help guide audio content towards the target loudness, but they are less "aggressive" and are only used when the audio content starts to deviate from the required target loudness.
[0044] In view of the above, the methods and apparatus described herein aim to make real-time processing conditions, also known as dynamic processing conditions, metadata-driven. Metadata enables dynamic loudness adjustment and dynamic range compression in real-time conditions. The methods and apparatus described herein have the advantage of: ● Real-time loudness adjustment and DRC use in MPEG-D DRC syntax; ● Real-time loudness adjustment and DRC use in combination with downmixId; ● Real-time loudness adjustment and DRC use in combination with drcSetId; ● Real-time loudness adjustment and DRC use in combination with eqSetId This makes it possible.
[0045] In other words, depending on the decoder settings (e.g., DRCSet, EQSet, and downmix), the decoder can look up a given payload for an appropriate pair of parameters and identifiers based on syntax by matching the above settings with identifiers. The parameters included in the pair whose identifiers best match the settings may then be selected as processing parameters for dynamic loudness adjustment to be applied to the original audio data received for correction.
[0046] Furthermore, multiple sets of parameters for dynamic processing (multiple instances of dynLoudCompValue) may be transmitted.
[0047] Metadata-driven dynamic loudness compensation can be used not only to correct overall loudness, but also to "center" the calculation and application of DRC gain. This centering can be the result of content loudness correction by dynamic loudness compensation and the way DRC is normally calculated and applied. In this sense, metadata for dynamic loudness compensation can be said to be used to adjust DRC parameters.
[0048] [Dynamic processing of audio data metadata databases] Referring to the example in Figure 1, a decoder 100 for dynamic processing of a metadata database of audio data for playback is described. The decoder 100 has one or more processors and non-temporary memory and may be configured to perform a method including the process shown in the example in Figure 2 by steps S101 to S105.
[0049] The decoder 100 may receive a bitstream containing audio data and metadata, and may, depending on the requirements, output raw (original) audio data, processed audio data after applying dynamic processing parameters determined from the metadata, and / or the metadata itself.
[0050] Referring to the example in Figure 2, in step S101, the decoder 100 may receive a bitstream containing audio data and metadata for dynamic loudness adjustment and optionally dynamic range compression (DRC). The audio data may be encoded audio data, or it may be unprocessed. That is, the audio data can be said to be original audio data. The metadata may contain multiple sets of parameters. For example, each payload of metadata may contain such multiple sets of metadata. These different sets of metadata may relate to each playback condition (e.g., different playback conditions).
[0051] The bitstream format is not restricted, but in embodiments, the bitstream may be an MPEG-D DRC bitstream. The presence of metadata for dynamic processing of audio data can then be signaled based on the MPEG-D DRC bitstream syntax. In embodiments, the loudnessInfoSetExtension() element may be used to carry the metadata as a payload, as further detailed below.
[0052] In step S102, the audio data and metadata can then be decoded by a decoder to obtain decoded audio data and metadata. In embodiments, the metadata may include one or more processing parameters relating to the average loudness value and optionally one or more processing parameters relating to the dynamic range compression characteristics. It is understood that each set of metadata may include its respective processing parameters.
[0053] Metadata allows for the application of dynamic or real-time corrections. For example, in the case of encoding and decoding for live real-time playback, the application of “real-time” or dynamic loudness metadata is desired to ensure that the live playback audio is properly loudness-controlled.
[0054] In step S103, the decoder then determines one or more processing parameters for dynamic loudness adjustment based on the playback conditions from the metadata. This may be done by using the playback conditions or information derived from the playback conditions (e.g., playback conditions information) to identify the appropriate set of metadata from among multiple sets of metadata.
[0055] In embodiments, the playback conditions may include one or more of the following: the device type of the decoder, the characteristics of the playback device, the characteristics of the loudspeaker, the loudspeaker setup, the characteristics of the background noise, the characteristics of the ambient noise, and the characteristics of the acoustic environment. Preferably, the playback condition information may indicate a specific loudspeaker setup. By considering the playback conditions, the decoder can select processing parameters for dynamic loudness adjustment in accordance with device and environmental constraints.
[0056] In an embodiment, the process of determining one or more processing parameters in step S130 may further include the decoder selecting at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix that corresponds to the playback conditions. Thus, at least one of the DRCSet, EQSet, and downmix correlates with or represents the constraints of individual devices and environments due to the playback conditions.
[0057] Preferably, step S103 includes selecting a set of DRC sequences (DRCSet). In other words, the selected set of metadata may include such a set of DRC sequences.
[0058] In an embodiment, the process determined in step S103 may further include identifying metadata identifiers that indicate at least one selected DRCSet, EQSet, and DownmixSet in order to determine one or more processing parameters from the metadata. The metadata identifiers thus enable the association of the metadata with the corresponding selected DRCSet, EQSet, and / or downmix, and thus with each playback condition.
[0059] In some embodiments, a specific loudspeaker setup may be used to determine the downmix, which can then be used to identify and select the appropriate one from a set of metadata. In such cases, the specific loudspeaker setup and / or downmix may be indicated by the playback conditions information described above.
[0060] In embodiments, the metadata may include one or more metadata payloads (e.g., dynLoudComp() payloads as shown in Table 5 below), each metadata payload may include multiple sets of parameters (e.g., the parameter dynLoudCompValue) and identifiers, each set including at least one of the DRCSet identifier (drcSetId), EQSet identifier (eqSetId), and downmix identifier (downmixId) in combination with one or more processing parameters relating to the identifier within the set. That is, each payload may have an array of entries, each entry including processing parameters and identifiers (e.g., drcSetId, eqSetId, downmixId). The array of entries may correspond to the above multiple sets of metadata. Preferably, each entry may have a downmix identifier.
[0061] In a further embodiment, the determination in step S103 may include selecting a set from among several sets in the payload based on the downmix selected by the decoder (or alternatively, based on at least one DRCSet, EQSet, and downmix), and the one or more processing parameters determined in step S103 may be one or more processing parameters related to identifiers in the selected set. That is, depending on the settings present in the decoder (e.g., DRCSet, EQSet, and downmix), the decoder can find a given payload for a suitable set of parameters and identifiers by matching the above settings with identifiers. The parameters included in the set in which the identifiers best match the settings may then be selected as processing parameters for dynamic loudness adjustment to be applied to the original audio data received for correction.
[0062] In step S104, one or more processing parameters determined may then be applied to the decoded audio data by the decoder to obtain processed audio data. The processed audio data, for example, live real-time audio data, is thus properly loudness-controlled.
[0063] In step S105, the processed audio data can then be output for playback.
[0064] In some embodiments, the bitstream may further include additional metadata for static loudness adjustment applied to the decoded audio data. Static loudness adjustment refers to the process performed for general loudness normalization, in contrast to dynamic processing for real-time situations.
[0065] By carrying metadata for dynamic processing separately from the additional metadata used for general loudness normalization, it becomes unnecessary to apply "real-time" corrections.
[0066] For example, in the case of encoding and decoding for live, real-time playback, the application of dynamic processing is desired to ensure that the live playback audio is properly loudness-controlled. However, in the case of transcoding for non-real-time playback, or where dynamic correction is not desired or unnecessary, the dynamic processing parameters determined from the metadata do not need to be applied.
[0067] By maintaining (dynamic / real-time) metadata for dynamic processing separately from additional metadata, the original, unprocessed content can be preserved as needed. The original audio is encoded along with the metadata. This allows playback devices to selectively apply dynamic processing, and furthermore, enables playback of the original audio content on high-end devices capable of playing the original audio.
[0068] As mentioned above, there are several advantages to maintaining dynamic loudness metadata separately from long-term loudness measurements / information such as contentLoudness (in ISO / IEC 23003-4). When combined, the content loudness (or what it should be after dynamic loudness metadata is applied) does not represent the actual loudness of the content, as the available metadata becomes a composite value. In addition to removing this ambiguity about what content loudness (or program or anchor loudness) is, there are several cases where this is particularly beneficial.
[0069] By separating metadata for dynamic processing, decoders or playback devices can turn off the application of dynamic processing and instead apply the implemented real-time loudness leveler, thus avoiding cascading leveling. This situation can occur, for example, if the device's own real-time leveling solution is superior to that used by the audio codec, or if the device's own real-time leveling solution is always active because it cannot be disabled, resulting in further processing that degrades resolution and impairs the playback experience.
[0070] By separating metadata for dynamic processing, transcoding to codecs that do not support dynamic loudness processing becomes possible, and custom loudness processing can be applied before re-encoding.
[0071] A further example is a live broadcast with a single encoding for the live feed. Dynamic processing metadata may be used or stored for archiving or on-demand services. Thus, in the case of archiving or on-demand services, more accurate or compliant loudness measurements can be performed based on the entire program, and the appropriate metadata can be reset.
[0072] In use cases where a fixed target loudness is used throughout the entire workflow, this is also beneficial, for example, in R128 compliance situations where -23LKFS is recommended. In this scenario, adding dynamic processing metadata is a “safety” measure, the content is expected and close to the required target, and adding dynamic processing metadata is a secondary check. Therefore, it is desirable to have the ability to turn it off. The content is expected and close to the required target, and adding dynamic processing metadata is a secondary check. Therefore, it is desirable to have the ability to turn it off.
[0073] [Encoding of audio data and metadata for dynamic loudness adjustment] Referring to the examples in Figures 3 and 4, an encoder is described that encodes the original audio data, dynamic loudness adjustments, and optionally metadata for dynamic range compression (DRC) into a bitstream, the encoder having one or more processors and non-temporary memory, and may be configured to perform a method including the process represented in the steps of the example in Figure 4.
[0074] In step S201, the original audio data may be input to the loudness leveler 201 for loudness processing so that the loudness-processed audio data is output from the loudness leveler 201.
[0075] In step S202, metadata for dynamic loudness adjustment may then be generated based on the loudness-processed audio data and the original audio data. Appropriate smoothing and time frames may be used to reduce artifacts.
[0076] In an embodiment, step S202 may include comparing the loudness-processed audio data with the original audio data by the analyzer 202. The metadata thus generated can emulate the leveling effect at the decoder site. The metadata is: ● Gain (wideband and / or multiband) processing parameters that, when applied to the original audio, generate loudness-compliant audio for playback; ● Processing parameters that indicate audio dynamics, for example ○ Peak-sample and true peak ○Short-term loudness value ○ Changes in short-term loudness values It may include.
[0077] In embodiments, step S202 may further include measuring loudness over one or more predefined periods using analyzer 202, and metadata may be generated based on the measured loudness. In embodiments, the measurement may include measuring the overall loudness of the audio data. Alternatively, or additionally, in embodiments, the measurement may include measuring the loudness of dialogue within the audio data.
[0078] In step S203, the original audio data and metadata can then be encoded into a bitstream. The format of the bitstream is not limited, but in embodiments, the bitstream may be an MPEG-D DRC bitstream, in which case the presence of metadata can be signaled based on the MPEG-D DRC bitstream syntax. In this case, in embodiments, a loudnessInfoSetExtension() element may be used to carry the metadata as a payload, as will be further detailed below.
[0079] In embodiments, metadata may include one or more metadata payloads, each metadata payload may include multiple sets of parameters and identifiers, each set including at least one of DRCSet identifiers (drcSetId), EQSet identifiers (eqSetId), and downmix identifiers (downmixId) in combination with one or more processing parameters relating to the identifiers within the set, the one or more processing parameters may be parameters for dynamic loudness adjustment by the decoder. In this embodiment, at least one of drcSetId, eqSetId, and downmixId may relate to at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix selected by the decoder. Generally, metadata can be said to include multiple sets of metadata, each set corresponding to different playback conditions (e.g., different playback conditions).
[0080] In some embodiments, the method may further include generating additional metadata for static loudness adjustment used by the decoder. Separating the metadata for dynamic loudness processing from the additional metadata within the bitstream, and further encoding the original audio data within the bitstream, offers several advantages, as detailed above.
[0081] The methods described herein may be carried out by a decoder or an encoder, respectively, which may have one or more processors and non-temporary memory and be configured to perform the above methods. An example of a device having such processing capability is shown in Figure 5, which shows a device 300 including two processors 301 and non-temporary memory 302.
[0082] Furthermore, the methods described herein may be further implemented in a system having an encoder for encoding metadata for original audio data, dynamic loudness adjustment, and optionally for dynamic range compression (DRC) into a bitstream, and a decoder for dynamic processing of the metadata database of the audio data for playback.
[0083] The method may be further implemented as a computer program product having a computer-readable storage medium containing instructions adapted to cause a device with processing capabilities to perform the method described above. The computer program product may be stored on the computer-readable storage medium.
[0084] [MPEG-D DRC Modified Bitstream Syntax] The following describes how the MPEG-D DRC bitstream syntax described in ISO / IEC 23003-4 may be modified in accordance with the embodiments described herein.
[0085] The MPEG-D DRC syntax can also carry dynamic processing metadata as frame-based dynLoudComp updates, as can be extended with the loudnessInfoSetExtension() element shown in Table 2 below.
[0086] For example, another switch case, UNIDRCLOUDEXT_DYNLOUDCOMP, may be added to the loudnessInfoSetExtension() element, as shown in Table 1. The switch case UNIDRCLOUDEXT_DYNLOUDCOMP may be used to identify the new element dynLoudComp(), as shown in Table 5. The loudnessInfoSetExtension() element may be an extension of the loudnessInfoSet() element, as shown in Table 2. Furthermore, the loudnessInfoSet() element may be a part of the uniDRC() element, as shown in Table 3. [Table 1] [Table 2] [Table 3] [Table 4] New dynLoudComp(): [Table 5] ● drcSetId allows dynLoudComp (regarding metadata) to be applied to each DRC set. ●eqSetId allows dynLoudComp to be applied in combination with various settings of the equalization tool. ●downmixId allows dynLoudComp to be applied to each DownmixId.
[0087] In some cases, in addition to the parameters described above, it may be beneficial to include a methodDefinition parameter (e.g., specified by 4 bits) that specifies the loudness measurement method used to derive dynamic program loudness metadata (e.g., anchor loudness, program loudness, short-term parameters, momentary loudness, etc.) and / or a measurementSystem parameter (e.g., specified by 4 bits) that specifies the loudness measurement system used to measure dynamic program loudness metadata (e.g., EBU R.128, ITU-R BS-1770 with or without preprocessing, ITU-R BS-1771, etc.). Such parameters may be included, for example, between the downmixId parameter and the dynLoudCompValue parameter within the dynLoudComp() element.
[0088] [Alternative syntax 1] [Table 6] [Table 7] [Table 8] In some cases, it may be beneficial to modify the syntax shown above in Table 8 so that the dynLoudCompPresent parameter and (when dynLoudCompPresent==1) the dynLoudCompValue parameter follow the reliability parameter inside the measurementCount loop of loudnessInfoV2(), rather than being outside the measurementCount loop. Furthermore, it may be beneficial to set dynLoudCompValue to equal to 0 when dynLoudCompPresent is 0.
[0089] [Alternative Syntax 2] Alternatively, the dynLoudComp() element may be placed in uniDrcGainExtension(). [Table 9] [Table 10] [Table 11] Semantics dynLoudCompValue: This field contains the value of dynLoudCompDb. The value is encoded according to the table below. The default value is 0dB. [Table 12] [Updated MPEG-D DRC loudness normalization processing] [Table 13] [Pseudocode for selecting and processing dynLoudComp]
number
number
number
number
[0090] [Alternative updated MPEG-D DRC loudness normalization process] [Table 14] When the alternative loudness normalization processes in Table 14 above are used, the loudness normalization pseudocode described above may be replaced by the following alternative loudness normalization pseudocode. Note that a default value for dynLoudCompDb, for example 0dB, may be assumed to ensure that a value for dynLoudCompDb is defined even when dynamic loudness metadata is not present in the bitstream.
number
[0091] targetLoudness: This field contains the desired output loudness. The value is encoded according to the following table. [Table 22] dynLoudnessNormalizationOn: This flag indicates whether dynamic loudness normalization processing is turned on or off. The default value is 0. If dynLoudnessNormalizationOn==0, dynloudnessNormalizationGainDb should be set to 0dB.
[0092] [interpretation] Unless otherwise stated, as will be evident from the following discussion, throughout this disclosure, discussions using terms such as “processing,” “computing,” “determining,” and “analyzing” will be understood to refer to the operation and / or processing of a computer or computing system, or similar electronic computing device, that manipulates and / or converts data, which is expressed as physical quantities such as electrons, into other data, which is similarly expressed as physical quantities.
[0093] Similarly, the term “processor” may refer to any device or part of a device that processes electronic data from, for example, registers and / or memory and converts that electronic data into other electronic data that can be stored, for example, in registers and / or memory. A “computer,” “computing machine,” or “computing platform” may include one or more processors.
[0094] The method logic described herein is executable by one or more processors that accept computer-readable (also called machine-readable) code, which includes a set of instructions that perform at least one of the methods described herein, when executed by one or more processors in an exemplary embodiment. This includes any processor capable of executing a set of instructions (sequential or otherwise) that specify an action to be performed. Thus, an example is a typical processing system comprising one or more processors. Each processor may include one or more of the following: a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem, which includes main RAM and / or static RAM, and / or ROM. A bus subsystem may be included for communication between components. The processing system may further be a distributed processing system, which includes processors connected by a network. If the processing system requires a display, such a display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. Where manual data entry is required, the processing system also includes input devices such as one or more alphanumeric input units such as keyboards, and instruction / control devices such as mice. The processing system may also include a storage system such as a disk drive unit. In some configurations, the processing system may also include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that carries computer-readable code (e.g., software) containing sets of instructions that cause one or more executions of the methods described herein when executed by one or more processors. Where the method includes several elements, for example, several steps, the order of such elements is not implied unless otherwise specified. The software may reside on a hard disk, or it may reside entirely or at least partially in RAM and / or in a processor during its execution by the computer system.Therefore, memory and processors also constitute a computer-readable carrier medium that carries computer-readable code. Furthermore, the computer-readable carrier medium may form or be included in a computer program product.
[0095] In alternative, exemplary embodiments, a higher processor may operate as a standalone device or be connected to other processors, for example, networked, in a networked configuration, one or more processors may operate as a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular telephone, web appliance, network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify the actions to be performed by the machine.
[0096] Furthermore, the term “machine” should also be interpreted to include any set of machines that individually or collectively execute a set (or set) of instructions to perform one or more of the method logics discussed herein.
[0097] Therefore, each example embodiment of the methods described herein takes the form of a computer-readable carrier medium that carries a set of instructions, for example, a computer program that runs on one or more processors, for example, one or more processors that are part of a web server deployment. Thus, as will be understood by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a special-purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, for example, a computer program product. The computer-readable carrier medium carries computer-readable code that, when executed on one or more processors, causes one or more processors to implement the method. However, embodiments of the present disclosure may take the form of a method, an exemplary hardware embodiment as a whole, an exemplary software embodiment as a whole, or an exemplary embodiment that combines embodiments of software and hardware. Furthermore, the present disclosure may take the form of a carrier medium (for example, a computer program product on a computer-readable storage medium) that carries computer-readable program code embodied in the medium.
[0098] The software may further be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in exemplary embodiments, the term “carrier medium” should be interpreted to include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term “carrier medium” should also be interpreted to include any medium capable of storing, encoding, or carrying sets of instructions executed by one or more processors, causing one or more processors to execute one or more of the method logics of this disclosure. The carrier medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wiring, and optical fibers, including wiring with bus subsystems. Transmission media can also take the form of sound waves or light waves, such as those generated during radio and infrared data communications. For example, the term “carrier medium” should be understood to include, but not limited to, computer products embodied in solid-state memory, optical media and magnetic media, media detectable by at least one processor or more processors that carry propagating signals representing sets of instructions that, when executed, perform a method, and transmission media in a network that are detectable by at least one of more processors and carry propagating signals representing sets of instructions.
[0099] It will be understood that, in an exemplary embodiment, the steps of the method discussed are performed by a suitable processor(s) of a processing (e.g., computer) system that executes instructions (computer-readable code) stored in storage. It will also be understood that this disclosure is not limited to any particular implementation or programming technique, and may be implemented using any suitable technique for implementing the functionality described herein. This disclosure is not limited to any particular programming language or operating system.
[0100] Any reference throughout this disclosure to “one embodiment,” “several embodiments,” or “exemplary embodiments” means that any particular feature, structure, or characteristic described in relation to an embodiment is included in at least one embodiment of this disclosure. Thus, the appearance of the phrases “in one embodiment,” “several embodiments,” or “exemplary embodiments” in various places throughout this disclosure does not necessarily all refer to the same exemplary embodiment. Furthermore, any particular feature, structure, or characteristic may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.
[0101] Unless otherwise specified as used herein, the use of ordinal adjectives such as “first,” “second,” and “third” to describe a common subject is merely to indicate that different instances of the same subject are being referred to, and is not intended to imply that the subjects described in this manner must be in a given order in time, space, rank, or otherwise.
[0102] In the following claims and descriptions herein, any one of the terms “comprising,” “comprised of,” or “which comprises” is an open term meaning that it includes, but does not exclude, at least the following elements / features. Therefore, when used in the claims, the term “comprising” should not be interpreted as limiting it to the means, elements, or steps listed below. For example, the expression “a device comprising A and B” should not be limited to a device consisting only of elements A and B. Any one of the terms “including” or “which includes” or “that includes” as used herein is also an open term meaning that it includes, but does not exclude, at least the following elements / features. Therefore, “including” is synonymous with “comprising.”
[0103] In the above description of exemplary embodiments of the Disclosure, it should be understood that various features of the Disclosure are sometimes grouped together in a single exemplary embodiment, figure, or description thereof in order to simplify the Disclosure and aid in understanding one or more of the various aspects of the Invention. However, this method of the Disclosure should not be construed as reflecting an intention that the claims require more features than are expressly described in each claim. Rather, as reflected in the subsequent claims, the aspects of the Invention lie in features that are not all present in a single exemplary embodiment disclosed above. Thus, the claims following the specification are incorporated herein, and each claim stands independently as a separate exemplary embodiment of the Disclosure.
[0104] Furthermore, while some exemplary embodiments described herein include some features included in other exemplary embodiments but not others, combinations of features of different exemplary embodiments are intended to fall within the scope of this disclosure and will be understood by those skilled in the art to form different exemplary embodiments. For example, any of the claimed exemplary embodiments can be used in any combination within the following claims.
[0105] Numerous specific details are provided in the description herein. However, it should be understood that exemplary embodiments of this disclosure may be carried out without referring to these specific details. In other instances, well-known methods, structures, and techniques are not described in detail so as not to obscure the understanding of this specification.
[0106] Therefore, while what is described here is believed to be the best mode of this disclosure, those skilled in the art will recognize that other and further modifications may be made to it without departing from the spirit of this disclosure, and it is intended that all such modifications and variations be claimed to be within the scope of this disclosure. For example, any formula given above merely represents a procedure that may be used. Functions may be added to or removed from the block diagram, and actions may be interchanged between function blocks. Steps may be added to or removed in the manner described within the scope of this disclosure.
[0107] The following enumerate example embodiments (EEE) describe some structural, feature, and function aspects of some aspects of the exemplary embodiments disclosed herein.
[0108] EEE1. A method for dynamic processing of a metadata database of audio data for playback, (a) The decoder receives a bitstream containing audio data and metadata for dynamic loudness adjustment, (b) Decoding the audio data and metadata to obtain the decoded audio data and metadata using the decoder, (c) The decoder determines one or more processing parameters for dynamic loudness adjustment based on playback condition information, (d) Applying one or more processing parameters determined above to the decoded audio data to obtain processed audio data, (e) Outputting the processed audio data for playback A method of having.
[0109] Eez2. The metadata indicates processing parameters for dynamic loudness adjustment for multiple playback conditions. Methods described in EEE1.
[0110] EEE3. Determining one or more processing parameters further includes determining one or more processing parameters for dynamic range compression (DRC) based on the playback conditions. The method described in EEE1 or EEE2.
[0111] EEE4. The playback conditions include one or more of the following: the device type of the decoder, the characteristics of the playback device, the characteristics of the loudspeaker, the loudspeaker setup, the characteristics of background noise, the characteristics of ambient noise, and the characteristics of the acoustic environment. The method described in any one of EEE1 to EEE3.
[0112] EEE5. Process (c) further includes the decoder selecting one of the following: a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix, which corresponds to the playback conditions. The method described in any one of EEE1 to EEE4.
[0113] EEE6. Process (c) further includes identifying metadata identifiers indicating the at least one selected DRCSet, EQSet, and downmix in order to determine the one or more processing parameters from the metadata, Methods for EEE5.
[0114] EEE7. The metadata includes one or more processing parameters relating to the average loudness value, and optionally one or more processing parameters relating to the dynamic range compression characteristics. The method described in any one of EEE1 to EEE6.
[0115] Eee8. The bitstream further includes additional metadata for static noise adjustment applied to the decoded audio data. A method according to any one of EEE1 to EEE7.
[0116] EEE9. The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The method described in any one of EEE1 through EEE8.
[0117] EEE10. The loudnessInfoSetExtension() element is used to carry the aforementioned metadata as a payload. Methods used in EEE9.
[0118] EEE11. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including at least one of the DRCSet identifier (drcSetId), EQSet identifier (eqSetId), and downmix identifier (downmixId) in combination with one or more processing parameters relating to the identifiers within the set. The method described in any one of EEE1 to EEE10.
[0119] EEE12. Process (c) includes selecting a set from among several sets in the payload based on at least one DRCSet, EQSet, and downmix selected by the decoder, The one or more processing parameters determined in process (c) are one or more processing parameters related to the identifiers in the selected set. Methods of EEE11 that are subordinate to EEE5.
[0120] EEE13. A decoder for dynamic processing of a metadata database of audio data for playback, the decoder comprising one or more processors and non-temporary memory, (a) The decoder receives a bitstream containing audio data and metadata for dynamic loudness adjustment, (b) Decoding the audio data and metadata to obtain the decoded audio data and metadata using the decoder, (c) The decoder determines one or more processing parameters for dynamic loudness adjustment based on playback condition information, (d) Applying one or more processing parameters determined above to the decoded audio data to obtain processed audio data, (e) Outputting the processed audio data for playback A decoder configured to perform a method having the following characteristics.
[0121] EEE14. A method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream, (a) Inputting the original audio data into the loudness leveler for loudness processing so that the loudness-processed audio data is obtained as the output from the loudness leveler, (b) Generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data, (c) Encoding the original audio data and the metadata into the bitstream A method of having.
[0122] EEE15. It further has the ability to generate additional metadata for static crowdness adjustment used by the decoder. Methods described in EEE14.
[0123] EEE16. Process (b) includes comparing the loudness-processed audio data with the original audio data, wherein the metadata is generated based on the results of the comparison. The method described in EEE14 or EEE15.
[0124] EEE17. Process (b) further includes measuring loudness over one or more predefined periods, wherein the metadata is further generated based on the measured loudness. The method described in EEE16.
[0125] EEE18. The measurement described above involves measuring the overall loudness of the audio data. Methods for EEE17.
[0126] EEE19. The measurement described above involves measuring the loudness of the dialogue in the audio data. Methods for EEE17.
[0127] EEE20. The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The method described in any one of EEE14 to EEE19.
[0128] EEE21. The loudnessInfoSetExtension() element is used to carry the aforementioned metadata as a payload. Methods used in EEE20.
[0129] Eez22. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including at least one of the DRCSet identifier (drcSetId), EQSet identifier (eqSetId), and downmix identifier (downmixId) in combination with one or more processing parameters relating to the identifier within the set, the one or more processing parameters being parameters for dynamic loudness adjustment by the decoder. The method described in any one of EEE14 to EEE21.
[0130] EEE23. At least one of drcSetId, eqSetId, and downmixId relates to at least one of a set of DRC sequences (DRCSet), a set of equalizer parameters (EQSet), and a downmix selected by the decoder. Methods described in EEE22.
[0131] EEE24. An encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, the encoder comprising one or more processors and non-temporary memory, (a) Inputting the original audio data into the loudness leveler for loudness processing so that the loudness-processed audio data is obtained as the output from the loudness leveler, (b) Generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data, (c) Encoding the original audio data and the metadata into the bitstream An encoder configured to perform a method having
[0132] EEE25. An encoder according to claim 24 for encoding original audio data and metadata for dynamic loudness adjustment and dynamic range compression (DRC) into a bitstream, A decoder for dynamic processing of the metadata database for playback, as described in EEE13. A system that has
[0133] EEE26. A computer program product having a computer-readable medium that includes instructions adapted to cause a device having processing capabilities to perform any one of the methods described in EEE1 to EEE12 or EEE14 to EEE23 when executed by said device.
[0134] EEE27. A computer-readable storage medium that stores computer program products as described in EEE26.
[0135] Eee28. The decoder receives, via the interface, an instruction on whether or not to perform the metadata database dynamic processing on the audio data for playback. When the decoder receives an instruction not to perform the metadata dynamic processing of the audio data for playback, it bypasses the step of applying at least one of the determined processing parameters to the decoded audio data. A method according to any one of EEE1 to EEE12, further comprising the above.
[0136] EEE29. Until the decoder receives the instruction via the interface whether to perform the metadata dynamic processing of the audio data for playback, the decoder bypasses the step of applying at least one of the determined processing parameters to the decoded audio data. Method described in EEE28.
[0137] EEE30. The metadata indicates multiple processing parameters for dynamic loudness adjustment for multiple playback conditions, The metadata further includes parameters that specify a loudness measurement method used to derive processing parameters among the plurality of processing parameters, The method described in any one of EEE1 to EEE12, EEE28, or EEE29.
[0138] EEE31. The metadata indicates multiple processing parameters for dynamic loudness adjustment for multiple playback conditions, The metadata further includes parameters that specify a loudness measurement system used to measure the processing parameters among the plurality of processing parameters. The method described in any one of EEE1 to EEE12 or EEE28 to EEE30.
[0139] [Cross-references to related applications] This application claims priority to European Patent Application No. 21193209.0 filed on 26 August 2021, U.S. Provisional Patent Application No. 63 / 237231 filed on 26 August 2021, and U.S. Provisional Patent Application No. 63 / 251307 filed on 1 October 2021. All of these applications are incorporated herein by reference in their entirety.
Claims
1. A method for dynamic processing of a metadata database of audio data for playback, The decoder receives a bitstream containing audio data and metadata for dynamic loudness adjustment, wherein the metadata for dynamic loudness adjustment includes multiple sets of metadata, each set of metadata corresponding to a different playback condition. The decoder decodes the audio data and metadata to obtain the decoded audio data and metadata, In response to the playback condition information supplied to the decoder, a set of metadata corresponding to a specific playback condition is selected, and one or more processing parameters for dynamic loudness adjustment are extracted from the selected set of metadata. Applying one or more of the extracted processing parameters to the decoded audio data to obtain processed audio data, The processed audio data is output for playback. It has, The selected metadata set includes a set of dynamic range compression (DRC) sequences (DRCSet), The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including each downmix identifier (downmixId) in combination with one or more processing parameters relating to the downmix identifier within that set. method.
2. The loudnessInfoSetExtension() element is used to carry the aforementioned metadata as a payload. The method according to claim 1.
3. A decoder for dynamic processing of a metadata database of audio data for playback, the decoder comprising one or more processors and non-temporary memory, The one or more processors described above are: The decoder receives a bitstream containing audio data and metadata for dynamic loudness adjustment, wherein the metadata for dynamic loudness adjustment includes multiple sets of metadata, and each set of metadata corresponds to a different playback condition. The decoder decodes the audio data and metadata to obtain the decoded audio data and metadata, In response to the playback condition information supplied to the decoder, a set of metadata corresponding to a specific playback condition is selected, and one or more processing parameters for dynamic loudness adjustment are extracted from the selected set of metadata. Applying one or more of the extracted processing parameters to the decoded audio data to obtain processed audio data, The processed audio data is output for playback. Configured to perform a method having, The selected metadata set includes a set of dynamic range compression (DRC) sequences (DRCSet), The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including each downmix identifier (downmixId) in combination with one or more processing parameters relating to the downmix identifier within that set. decoder.
4. A method for encoding audio data and metadata for dynamic loudness adjustment into a bitstream, The process involves inputting the original audio data into the loudness leveler for loudness processing so that the loudness-processed audio data is obtained as the output from the loudness leveler, The process involves generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data, Encoding the original audio data and the metadata into the bitstream. It has, The metadata includes multiple sets of metadata, each set of metadata corresponding to a specific playback condition. The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including a downmix identifier (downmixId) combined with one or more processing parameters relating to the downmix identifier within that set, and the one or more processing parameters are parameters for dynamic loudness adjustment by the decoder. method.
5. It further has the ability to generate additional metadata for static crowdness adjustment used by the decoder. The method according to claim 4.
6. Generating the metadata includes comparing the loudness-processed audio data with the original audio data, and the metadata is generated based on the results of the comparison. The method according to claim 4.
7. Generating the metadata further includes measuring loudness over one or more predefined periods, and the metadata is further generated based on the measured loudness. The method according to claim 6.
8. The measurement described above involves measuring the overall loudness of the audio data. The method according to claim 7.
9. The measurement described above involves measuring the loudness of the dialogue in the audio data. The method according to claim 8.
10. The loudnessInfoSetExtension() element is used to carry the aforementioned metadata as a payload. The method according to claim 4.
11. An encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, the encoder comprising one or more processors and non-temporary memory, The one or more processors described above are: The process involves inputting the original audio data into the loudness leveler for loudness processing so that the loudness-processed audio data is obtained as the output from the loudness leveler, The process involves generating metadata for dynamic loudness adjustment based on the loudness-processed audio data and the original audio data, Encoding the original audio data and the metadata into the bitstream. Configured to perform a method having, The metadata includes multiple sets of metadata, each set of metadata corresponding to a specific playback condition. The bitstream is an MPEG-D DRC bitstream, and the presence of the metadata is signaled based on the MPEG-D DRC bitstream syntax. The metadata includes one or more metadata payloads, each metadata payload includes multiple sets of parameters and identifiers, each set including a downmix identifier (downmixId) combined with one or more processing parameters relating to the downmix identifier within that set, and the one or more processing parameters are parameters for dynamic loudness adjustment by the decoder. Encoder.
12. An encoder for encoding original audio data and metadata for dynamic loudness adjustment into a bitstream, as described in claim 11, A decoder for dynamic processing of a metadata database of audio data for playback, as described in claim 3, A system that has
13. A computer program having instructions adapted to cause a device having processing capabilities to perform the method according to any one of claims 1 to 2 or 4 to 10, when executed by said device.
14. A computer-readable storage medium storing the computer program described in claim 13.
Citation Information
Patent Citations
Multi-channel audio signal encoding and decoding method and apparatus
JP2009512893A
Loudness adjustment for downmixed audio content
JP2016534669A
Encoded audio metadata-based loudness equalization and dynamic equalization during drc
JP2020008878A
Optimizing loudness and dynamic range across different playback devices
JP2021089444A