Loudness control for user interaction in an audio coding system
The audio processor addresses the challenge of inconsistent loudness in user-interacted audio systems by dynamically adjusting loudness based on user input and metadata, ensuring consistent playback levels.
Patent Information
- Application Number
- JP2025187139
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-03
AI Technical Summary
Existing audio coding systems lack effective means for loudness compensation that can adapt to user interactions, leading to inconsistent loudness levels during playback.
An audio processor that includes an audio signal modification unit, a loudness control unit, and a loudness manipulation unit to dynamically adjust loudness based on user input and metadata, ensuring consistent loudness compensation by determining a loudness compensation gain using reference and modified loudness values.
Enables consistent loudness normalization even with user interactions, preserving overall loudness levels and adapting to different playback configurations.
Smart Images

Figure 2026016762000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio processor and an audio encoder, and also to a corresponding method. [Background technology]
[0002] Modern audio coding systems not only provide a means for efficiently transmitting audio content in a loudspeaker-channel-based representation for simple playback at the decoder, but also include more advanced features that allow users to interact with the content and therefore influence how the audio is played and rendered at the decoder. This enables new forms of user experience compared to legacy audio coding systems.
[0003] An example of an advanced audio coding system is the MPEG-H 3D Audio standard (Non-Patent Document 1), which allows the transmission of immersive audio content in three different formats: channel-based, object-based, and scene-based using higher order ambisonics (HOA). It has been designed to offer new possibilities such as personalized user interaction and audio adaptation to different usage scenarios.
[0004] The three different categories of content formats can be described as follows: -Channel-based: Traditionally, spatial audio content (starting with simple two-channel stereo) has been delivered as a set of precisely defined channel signals directed to be reproduced by loudspeakers at fixed target positions relative to the listener. - Object-based: An audio object is a signal that should be played back to come from a specific target location, specified by associated side information provided as metadata with the audio. In contrast to channel signals, the actual placement of audio objects can change over time and does not need to be predefined during the sound production process, but is determined by rendering to a target loudspeaker setup at the time of playback. The actual placement of audio objects may include user interactivity regarding the location or level of an object or group of objects. Higher Order Ambisonics (HOA) is an alternative approach to capturing a 3D sound field by transmitting several "coefficient signals" that have no direct relationship to the channels or objects. The actual object signals for reproduction are generated at the decoder, taking into account a given loudspeaker configuration.
[0005] A method for loudness compensation in an object-based audio coding system including user interaction is disclosed in US Pat. No. 6,233,669. A decoder receives an audio input signal including audio object signals and generates an audio output signal. A signal processor determines loudness compensation values for the audio output signal based on loudness information associated with the audio input signal and rendering information. The rendering information indicates whether one or more audio object signals should be amplified or attenuated, and may be adjusted according to user preferences. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] European Patent Application Publication No. 2879131 [Non-patent literature]
[0007] [Non-Patent Document 1] J. Herre at al., “MPEG-H Audio - The New Standard for Universal Spatial / 3D Audio Coding”, 137th AES Convention, 2014, Los Angeles Summary of the Invention [Problem to be solved by the invention]
[0008] The object of the present invention is to improve the feasibility of loudness compensation. [Means for solving the problem]
[0009] The object is achieved by an audio processor for processing an audio signal, the audio processor comprising the following elements: an audio signal modification unit configured to modify the audio signal in response to a user input, a loudness control unit configured to determine a loudness-compensation gain based on a reference loudness or reference gain on the one hand and based on a modified loudness or modified gain on the other hand, the modified loudness or modified gain depending on the user input, the loudness control unit configured to determine the loudness-compensation gain based on metadata of the audio signal indicating which groups should or should not be used for determining the loudness-compensation gain, the groups comprising one or more audio elements, and a loudness manipulation unit configured to manipulate the loudness of a signal using the loudness-compensation gain.
[0010] An audio processor - or decoder or device for processing audio signals - receives an audio signal and in one embodiment generates an output signal, which includes audio objects, audio elements etc. of the audio signal to be played, for example by loudspeakers or earphones or to be stored on a medium etc.
[0011] The audio processor reacts to user input via an audio signal modification unit configured to modify the audio signal in response to the user input. The user input, in one embodiment, indicates amplification or attenuation of a group and / or switching off or on a group. The group may comprise one or more audio elements, such as audio objects, channels, objects, or HOA components. Depending on the embodiment, the user input may also refer to data related to a playback configuration used to play back the signal. A further user input may also refer to selection of a preset. A preset refers to a set of at least one group and, depending on the embodiment, may specify a uniquely measured group loudness value and / or gain value for each group. The user input is used by the audio modification unit to appropriately modify the audio signal. In one embodiment, the metadata includes data attributing to multiple presets.
[0012] In one embodiment, the preset refers to a set of one group, and in another embodiment, defines multiple groups that do not belong to the preset.
[0013] The audio processor further includes a loudness control unit configured to determine a loudness compensation gain. This loudness compensation gain, referred to herein as C, makes it possible to balance the effect of a user input in order to provide a signal with an overall loudness desired or set by the user. The loudness compensation gain is determined based on a reference loudness or reference gain, on the one hand, and a modified loudness or modified gain, on the other hand. Thus, the loudness compensation gain is determined based on the reference loudness or reference gain and the modified loudness or modified gain. The modified loudness or modified gain depends on the user input.
[0014] The loudness control unit is additionally configured to determine the loudness compensation gain based on metadata of the audio signal, the metadata associated with the audio signal conveying information about the audio signal and the individual groups and, in one embodiment, being included in the audio signal itself.
[0015] The metadata data (in embodiments of the audio processor described herein) indicates whether a group (particularly contained in the audio signal) should be used (e.g., considered) or not used (e.g., ignored) for determining the loudness compensation gain. Thus, information about the corresponding group is either considered or ignored for determining the loudness compensation gain. In at least one embodiment, whether one or more groups are considered or ignored further depends on user input.
[0016] In one embodiment, considering or ignoring a group includes partially considering or ignoring it in the sense that the group and its individual values are only used for determining part of the loudness compensation gain, e.g., only for calculating the reference loudness or the modified loudness.
[0017] The loudness compensation gain is used by a loudness manipulation unit included in an audio processor, which manipulates the loudness of a signal using the loudness compensation gain, and the loudness compensation gain applied is not only influenced by user input, but is also a result of metadata data associated with or belonging to the audio signal.
[0018] The signal manipulated by the loudness manipulation unit is, according to one embodiment, an output signal provided by an audio processor and based on the audio signal, wherein the loudness manipulation unit provides the output signal and manipulates the loudness of the output signal using a loudness compensation gain.
[0019] In a different embodiment, the loudness manipulation manipulates the loudness of a signal that is provided to the loudness manipulation and that is preferably already modified by user input, in which part of the audio processor provides or generates a signal that is fed to the loudness manipulation and processed accordingly, i.e. modified with respect to its loudness, by the loudness manipulation.
[0020] In a further embodiment, the signal whose loudness is manipulated by the loudness manipulation unit is an audio signal. In this case, the loudness manipulation unit modifies the metadata of the audio signal by the modification. This embodiment relates to a further embodiment in which an audio processor provides a modified audio signal. This modified audio signal has been modified according to user input and according to the loudness modification. This modified audio signal also later becomes a bitstream.
[0021] According to one embodiment of the audio processor, the loudness control unit is configured to determine the loudness compensation gain based on at least one flag included in the metadata data, which flag indicates whether or how a group should be considered for determining the loudness compensation gain. In this embodiment, the metadata includes, for example, a flag having a value of either "true" or "false" indicating whether the associated group should be considered for calculating the loudness compensation gain, respectively. Considering a group, in one embodiment, also refers to the question of which step of the calculation the group should be used for. This also refers, for example, to the calculation of the reference loudness and the modified loudness. The reference loudness and the modified loudness are the overall loudness calculated before and after the user input, respectively. In a different embodiment, the flag indicates that the corresponding group only exists for a short period of time and therefore can be ignored for determining the loudness compensation gain.
[0022] According to one embodiment of the audio processor, the loudness control unit is configured to use only some groups for determining the loudness compensation gain if these groups belong to anchors contained in the metadata of the audio signal, which anchors in one embodiment refer to audio elements belonging to, for example, voice, dialogue or special sound effects.
[0023] The handling of groups belonging to anchors will be further explained in the following embodiments.
[0024] In one embodiment, the loudness control unit is configured to use only groups belonging to an anchor for determining the loudness-compensation gain if the modified gain of at least one group belonging to said anchor is greater than the corresponding reference gain, so that if the gain value of at least one of these "anchor groups" is increased in response to user input, i.e. if the user has boosted at least one of these groups, only those groups of anchors will be used for calculating the loudness-compensation gain.
[0025] In an alternative or additional embodiment, the loudness control unit is configured to use the groups belonging to the anchor and the groups not belonging to the anchor to determine the loudness compensation gain when the modified gain of at least one group belonging to the anchor is smaller than the corresponding reference gain. Thus, in this embodiment, when the gain value of at least one anchor group is decreased due to a user input, not only the groups belonging to the anchor but also the groups not belonging to the anchor are used for the calculation.
[0026] In one embodiment, the two previous embodiments are combined: thus, modifying the gain of at least one group belonging to an anchor determines whether only the anchor group or the anchor group and non-anchor groups are used to determine the loudness compensation gain.
[0027] The object of the present invention is achieved by an audio processor for processing an audio signal, comprising the following elements: an audio signal modification unit configured to modify the audio signal in response to a user input, a loudness control unit configured to determine a loudness-compensation gain based on a reference loudness or reference gain on the one hand and based on a modified loudness or modified gain on the other hand, wherein the modified loudness or modified gain is dependent on the user input, and the loudness control unit configured to determine the loudness-compensation gain based on metadata of the audio signal indicating at least one preset, wherein the preset refers to a set of at least one group comprising one or more audio elements, and a loudness manipulation unit configured to manipulate the loudness of the signal using the loudness-compensation gain.
[0028] For a general description of audio processors, see the description above.
[0029] The loudness control unit of the audio processor references metadata data associated with or belonging to an audio signal. The data indicates a preset, where a preset refers to a set of at least one group containing one or more audio elements. This embodiment considers the case where a combination of groups is associated with a specific loudness and / or gain value for a specific preset. Therefore, the metadata includes data about groups depending on different presets or at least a default preset. Thus, the loudness control unit uses data associated with a preset selected by a user or that is the default preset.
[0030] The audio processor is in one embodiment configured according to at least one of the above-described embodiments, and thus the above-described embodiments are at least partially realized using the above-described audio processor.
[0031] According to an embodiment of the audio processor, the loudness control unit is configured to determine the loudness compensation gain based on the group loudness and / or gain values of at least one group of the set indicated by a preset, which preset refers to a specific set of groups of audio elements comprised in the audio signal, for which the metadata contains specific data, i.e., group loudness and / or gain values, to be used for determining the loudness compensation gain when the corresponding preset is selected or set as a default preset.
[0032] In a further embodiment, the loudness control unit is configured to determine a reference loudness for the set indicated by the preset using the individual group loudnesses and the individual gain values. The loudness control unit is also configured to determine a modified loudness for the set indicated by the preset using the individual group loudnesses and the individual modified gain values, the modified gain values being modified by user input. In this embodiment, the reference loudness and the modified loudness are determined based on values for groups associated with and belonging to a preset. The determination also carries instructions on whether and how the groups should be used—e.g., for determining the reference loudness or the modified loudness.
[0033] In a further embodiment, the loudness control unit is configured to determine the loudness compensation gain based on data included in the metadata of the audio signal indicating a selected preset, the preset being selected by user input, in this embodiment said preset being selected by a user via user input.
[0034] According to one embodiment of the audio processor, the loudness control unit is configured to determine the loudness compensation gains based on data included in the metadata of the audio signal indicating a default preset. This default preset may be set prior to or independently of user input. This embodiment handles the case where the user does not select a preset. To this end, the default preset may be used, for example, prior to any user input, thereby ensuring that a certain set of data, here covering the default preset, is used to determine the loudness compensation gains, even in the absence of, for example, user interaction.
[0035] The object of the present invention is achieved by an audio processor for processing an audio signal, comprising the following elements: an audio signal modification unit configured to modify the audio signal in response to a user input, a loudness control unit configured to determine a loudness compensation gain based on a reference loudness or reference gain on the one hand and based on a modified loudness or modified gain on the other hand, wherein the modified loudness or modified gain is dependent on the user input, the loudness control unit configured to determine the loudness compensation gain based on metadata of the audio signal indicating whether a group is switched off or switched on, wherein the group comprises one or more audio elements, and a loudness manipulation unit configured to manipulate the loudness of the signal using the loudness compensation gain.
[0036] For a general description of the audio processor of this embodiment, please see the description above.
[0037] Here, the loudness control unit is configured to determine the loudness compensation gain based on metadata of the audio signal indicating whether a group is switched off or on. In one example, the audio signal may contain, as audio objects, different soundtracks belonging to different language versions of a video. The presets may also indicate different language versions. Thus, for each different preset, one soundtrack for one language will be switched on while the other version will be switched off. This example also shows that a user can switch between different language versions and select the desired language version provided, thereby switching off the soundtrack associated with the default preset. However, switching on a group does not always mean switching off another group, and vice versa.
[0038] In one embodiment, the audio processor is configured according to at least one of the preceding embodiments.
[0039] Thus, the above-described embodiments may be at least partially implemented using the audio processor described above, which may also be implemented using the audio processor described below, in at least one embodiment contemplated by the embodiments described below.
[0040] According to one embodiment, the loudness control unit determines the loudness compensation gain based on a user input depending on whether a group is switched off or on by the user input, where the user interaction influences the loudness control unit gain determination.
[0041] According to one embodiment of the audio processor, the loudness control unit is configured to exclude a group for determining the modified loudness if said group is switched off in response to a user input, in this embodiment, when a user switches off a group, this group is not used for determining the modified loudness resulting from the loudness value representing the user preference.
[0042] In a further embodiment, the loudness control unit is configured to exclude a group for determining the reference loudness if the group is switched off in the metadata, and to include the group for determining the modified loudness if the group is switched on by user input. In this embodiment, a group is switched off in the metadata and is not used for determining the reference loudness. If the user switches on the group, the group is included for evaluating the modified loudness.
[0043] According to one embodiment of the audio processor, the loudness control unit is configured to include a group for determining the reference loudness when said group is switched on in the metadata, and to exclude said group for determining the modified loudness when said group is switched off by user input. In this embodiment, the opposite case to the previous embodiment is considered.
[0044] The object of the present invention is also achieved by an audio processor for processing an audio signal, comprising the following elements: an audio signal modification unit configured to modify the audio signal in response to a user input, a loudness control unit configured to determine a loudness compensation gain based on a reference loudness or reference gain on the one hand and on a modified loudness or modified gain on the other hand, wherein the modified loudness or the modified gain is dependent on the user input, the loudness control unit configured to determine the loudness compensation gain based on metadata of the audio signal missing at least one group loudness among a group of metadata contained in the audio signal, and a loudness manipulation unit configured to manipulate the loudness of the signal using the loudness compensation gain.
[0045] For a general description of the audio processor in this embodiment, please see the description above.
[0046] In this audio processor (or decoder), the loudness control section handles the situation where a corresponding group loudness is missing for a group present in the audio signal: the group loudness may be missing for a particular preset or playback configuration, etc., or the metadata may be completely empty of any group loudness for this group.
[0047] In one embodiment, the audio processor is configured according to at least one of the above-described embodiments. Thus, the above-described embodiment is at least partially realized using the audio processor described above. The above-described audio processor is likewise realized using the audio processor described below, in at least one embodiment taking into account the embodiments described below.
[0048] According to one embodiment of the audio processor, the loudness control unit is configured to calculate the missing group loudness using the preset loudness, the reference gains of the groups having the missing group loudness, and the group loudness and the reference gains for the groups having a group loudness, where the preset loudness is the overall loudness of the preset groups.
[0049] In a further embodiment, the loudness control unit is configured to determine a loudness compensation gain using only the at least one reference gain and the at least one modified gain for blind loudness compensation when the metadata of the audio signal lacks at least one group loudness. In this embodiment, the case where at least one group loudness is missing and the case where all group loudnesses are missing are treated in the same way.
[0050] According to one embodiment of the audio processor, the loudness control unit is configured to determine a loudness compensation gain using only at least one reference gain and at least one modified gain for blind loudness compensation when the metadata of the audio signal is invalid for group loudness.
[0051] The object of the present invention is also achieved by an audio processor for processing an audio signal, comprising the following elements: an audio signal modification unit configured to modify the audio signal in response to a user input, a loudness control unit configured to determine a loudness compensation gain based on a reference loudness or reference gain on the one hand and based on a modified loudness or modified gain on the other hand, wherein the modified loudness or modified gain is dependent on the user input, the loudness control unit configured to determine the loudness compensation gain based on metadata of the audio signal referring to a playback configuration for the reproduction of the signal, and a loudness manipulation unit configured to manipulate the loudness of the signal using the loudness compensation gain.
[0052] For a general description of the audio processor in this embodiment, please see the description above.
[0053] The audio processor determines the loudness compensation gain based on data indicative of a particular playback configuration. Metadata associated with the audio signal, and in one embodiment included in the audio signal, therefore includes data specific to at least one playback configuration. In one embodiment, for each playback configuration, the metadata includes data corresponding to the individual playback-or replay-configuration.
[0054] In one embodiment, the audio processor is configured according to at least one of the preceding embodiments, and thus in one embodiment, the audio processor is combined with at least one of the preceding embodiments.
[0055] According to one embodiment of the audio processor, the loudness control unit is configured to determine the loudness compensation gains based on metadata data indicating the playback configuration and including associated group loudness and / or reference gain values, such that different playback configurations are associated with different gain values and / or group loudness for the individual groups.
[0056] In one embodiment, the metadata includes data about different presets and different playback configurations.
[0057] In a further embodiment, the audio processor comprises a configuration conversion unit for converting data contained in the metadata and indicative of a playback configuration into data indicative of a current playback configuration, and the loudness control unit is configured to determine the loudness compensation gains using the data provided by the configuration conversion unit. In this embodiment, the audio processor handles situations where the current playback configuration for playback of the signal differs from the playback configuration provided by the metadata. Thus, the data of the metadata is converted to match the current playback configuration, and the converted data is used to determine the loudness compensation gains.
[0058] In one embodiment, the audio processor includes a format conversion unit for converting the signal into a predetermined playback configuration, hi a further embodiment the loudness control unit is configured to select a loudness value specific for the particular playback configuration used by the format conversion unit.
[0059] The following embodiments can be realized by any of the above-mentioned embodiments.
[0060] In one embodiment, the audio signal includes a bitstream having metadata, and the metadata includes a reference gain for at least one group.
[0061] According to one embodiment of the audio processor, the metadata of the audio signal comprises a group loudness for at least one group, hi a further embodiment, the metadata comprises a group loudness for multiple groups belonging to the audio signal.
[0062] In a further embodiment, the loudness control unit is configured to determine a reference loudness for at least one group using a group loudness and a gain value for—at least one—group, and the loudness control unit is configured to determine a modified loudness using the group loudness and a modified gain value, the modified gain value being modified by user input.
[0063] In one embodiment, the loudness control unit uses the individual group loudnesses, referred to as Li, and gain values, referred to as gi, of the plurality of groups to calculate a reference loudness, referred to as L, for the plurality of groups. ref Further, the loudness control unit is configured to determine a modified loudness L for the plurality of groups using the individual group loudnesses L and the modified gain values H. mod In one embodiment, the two groups are the same, and in another embodiment, they are different. The groups also depend on the individual data in the metadata.
[0064] In a further embodiment, the loudness control unit is configured to perform a limiting action on the loudness compensation gain so that the loudness compensation gain is below an upper threshold and / or so that the loudness compensation gain is above a lower threshold.
[0065] According to one embodiment of the audio processor, the loudness manipulation unit is configured to apply a corrected gain to the signal determined by a loudness compensation gain and a normalization gain, the normalization gain being determined by a target loudness level set by user input and a metadata loudness level included in the metadata of the audio signal. In one embodiment, the normalization gain is determined using a ratio between the loudness levels of the individual groups of the audio signal and the loudness level set by the user to be experienced by the user for playback of the audio signal.
[0066] The above-described embodiment of the audio processor allows loudness compensation according to user input, which is improved by taking into account data describing groups of audio signals and their associations or usage for loudness compensation.
[0067] The above embodiments refer to audio processors or audio decoders. In the following, an encoder is described, providing an audio signal associated with or including metadata to be used by the audio processor.
[0068] The object is achieved by an audio encoder for generating an audio signal including metadata, the audio encoder comprising a loudness determiner for determining a loudness value for at least one group including one or more audio elements, and a metadata writer for introducing said determined loudness value into said metadata as a group loudness.
[0069] According to one embodiment of the audio encoder, the loudness determiner is configured to determine different loudness and / or gain values for different playback configurations, and the metadata writer is configured to associate the determined different loudness and / or gain values with the respective playback configurations and introduce the determined different loudness and / or gain values into the metadata, in this embodiment the metadata contains different data for groups associated with the different playback configurations, thus improving the playback of each group of the audio signal.
[0070] In one embodiment, the loudness determiner is configured to determine different loudness values and / or different gain values for different presets that indicate a set of at least one group containing one or more audio elements. Further, the metadata writer is configured to associate the determined different loudness values and / or different gain values with the respective presets and introduce them into the metadata. In this embodiment, the presets indicate a particular set of groups that are associated with a particular group loudness and / or reference gain value.
[0071] In further embodiments, the audio encoder further comprises a controller configured to determine which groups should be used or ignored for determining the loudness compensation gain, and the metadata writer configured to write an indication into the metadata indicating which groups should be used or ignored for determining the loudness compensation gain. The indication is in one embodiment a flag. In some embodiments, the indication indicates a preset, a playback configuration, an anchor and / or a duration, thereby indicating the relevance of the groups.
[0072] In at least one embodiment, the metadata includes different data (eg, group loudness or reference gain) having different values for at least one group of audio signals.
[0073] According to one embodiment of the audio encoder, the audio encoder further comprises an estimator configured to calculate a group loudness value for a group, the group loudness value for the group not determined by the loudness determiner. The metadata writer is configured to introduce the calculated group loudness value into the metadata such that all groups of the audio signal have an associated group loudness. In this embodiment, the audio encoder compensates for the lost group loudness by calculating the group loudness based on valid data.
[0074] The object of the invention is also achieved by a method for processing an audio signal.
[0075] The method includes at least the following steps: modifying the audio signal in response to user input; determining a loudness compensation gain based on, on the one hand, a reference loudness (as the overall loudness of the relevant individual groups before being modified by the user) or a reference gain, and on the other hand, a modified loudness (as the combined loudness of the relevant groups after being modified by the user, paired with the reference loudness) or a modified gain, wherein the modified loudness or the modified gain depends on the user input. The determination of the loudness compensation gain - referred to as C - is performed using at least one or a combination of the following embodiments: the loudness compensation gain is determined based on metadata data associated with - or included in - the audio signal. In different embodiments, each group contains one or more audio elements, the data being: The data indicates whether certain groups contained in the audio signal should be taken into account or ignored to determine this loudness compensation gain. The data refers to a preset, which in turn refers to a set of at least one group. -The data shows whether a group is switched off or switched on. - In the data, the group loudness of at least one of the groups included in the audio signal is missing. The data refers to a playback configuration for reproducing the signal. Manipulating the loudness of the output signal relative to the audio signal using a loudness compensation gain.
[0076] The object of the present invention is also achieved by a method for generating an audio signal including metadata, the method comprising the steps of determining a loudness value for a group including one or more audio elements, and introducing the determined loudness value for the group into the metadata as a group loudness.
[0077] The objects of the invention are also achieved by a computer program for carrying out the above-mentioned method when running on a computer or processor.
[0078] An apparatus embodiment (whether an audio processor or an audio encoder) may also be performed by the steps of a method and corresponding embodiments of the method, and therefore the description of an apparatus embodiment is also applicable to the method.
[0079] The present invention will now be described with reference to the accompanying drawings, in which embodiments are shown. [Brief explanation of the drawings]
[0080] [Figure 1] FIG. 1 is a schematic diagram of an audio decoder. [Figure 2] 1 is a schematic diagram of an audio processor according to the present invention; [Figure 3] 1 is a schematic diagram of an inventive audio encoder; DETAILED DESCRIPTION OF THE INVENTION
[0081] FIG. 1 shows a schematic diagram of an MPEG-H 3D audio decoder as an example of an audio processor, showing all the main building blocks of the system. As a first step, the received audio stream 500 (containing the transmitted audio signal, which may be a channel, object or HOA component with associated metadata) is decoded by a decoder 501 to provide the audio content 502 and associated metadata 503. The channel signals are mapped to the target playback loudspeaker setup using a format converter 504, which acts as a channel renderer and format converter. · The object signals are rendered by the object renderer 505 using the associated object metadata to the target playback loudspeaker setup. Higher-order Ambisonics content is rendered by the HOA renderer 506 to the target playback loudspeaker configuration using the associated HOA metadata. The loudspeaker signals corresponding to the different elements (channels, objects, HOAs) in the form of an audio signal 507 as output of the format converter 504, the object renderer 505 and the HOA renderer 506 are then mixed together in a mixing stage. This is performed by a mixer 508, providing a mixed audio signal 509. The output 509 of the mixer 508 is then processed by a loudness control stage, where the audio signal is normalized to a desired target loudness level. A loudness control unit 510 performs loudness compensation in addition to the normalization. For this purpose, the loudness control unit 510 receives a user input 511. The user input 511, resulting from a user interaction, also refers to information about the loudness configuration to be used for playback, and is also provided to the format converter 504, the object renderer 505, and the HOA renderer 506. The loudness control unit 510 is supplied with metadata 503 extracted by the decoder 501 from the received audio stream 500, which in particular refers to rendering and / or loudness information. The resulting signal 512 is, in the illustrated embodiment, provided to the loudspeakers of the available loudspeaker configuration for playback.
[0082] The possible user interactivity can be divided into, for example, two different categories. -Selection of presets for transmitted audio programs Manipulating default rendering for groups of audio elements
[0083] The meaning of presets and groups in the context of MPEG-H 3D audio and the present invention is explained below.
[0084] The individual channels, objects, and HOA scenes available for a transmitted audio program are called audio elements. A group refers to a specific collection of individual audio elements. The specific grouping information for audio elements is contained in the MPEG-H 3D audio metadata, which is transmitted together with the audio content in the audio stream. Elements of a group cannot be independently and interactively modified; only the group as a whole can be manipulated, i.e., all contained elements are manipulated together. An example is given by a group consisting of channels corresponding to a stereo or 5.1 loudspeaker configuration. In special cases, a group can consist of only a single element, for example, a dialogue object of a program. In that case, the user can change the level of this dialogue object within the audio scene, for example.
[0085] A preset defines a combination of groups in an audio scene. Presets can be used to efficiently signal different representations of the same audio program within the same audio stream. This preset definition also contains default or initial rendering information for each group, which is used if no user action is applied. The most important example of this rendering information is the gain applied to a group when rendering the entire audio scene. The configuration information defining a preset is determined by the encoder and is part of the metadata, e.g., MPEG-H 3D audio metadata.
[0086] It should be noted that a primary or default audio scene can be thought of as a special form of preset that includes all audio elements without necessarily specifying grouping information, however default or initial rendering information (e.g. gain) for individual audio elements is typically also provided in the metadata for the primary audio scene.
[0087] One of the most important features for the next generation of audio distribution is advanced loudness control, i.e., proper signaling of loudness information and loudness normalization. Loudness control is particularly important in broadcast applications, where it represents an essential feature for satisfying applicable broadcasting regulations and recommendations.
[0088] The loudness control scheme included in MPEG-H 3D Audio is based on metadata representing the measured loudness of an audio program. The metadata is transmitted together with the actual audio content in an audio stream as an example of an audio signal to be processed by an audio processor. In a decoder according to one embodiment, a loudness normalization gain is calculated based on the transmitted loudness information and a target loudness level. In one embodiment, the loudness normalization gain is applied to the audio signal after the mixer 508, as shown by way of example in FIG. 1.
[0089] To realize the unique feature of having multiple presets of the same audio program within the same audio stream, additional loudness metadata corresponding to the measured loudness of each preset is included in the audio stream. Processing steps such as format conversion (downmixing) or dynamic range processing can potentially change the loudness of the audio. Therefore, in one embodiment, additional loudness information is included in these cases to ensure accurate loudness normalization.
[0090] In other embodiments, loudness information is transmitted for individual groups or even for single audio elements. In one embodiment, group loudness information is provided for different loudspeaker configurations. For example, if a group consists of multiple channel signals, different group loudspeaker information may be included assuming playback to a stereo or 5.1 loudspeaker configuration. The group loudness information may be used for loudness control in the interactive scenarios proposed in this invention.
[0091] The loudness information described above refers to a large variety of configurations for a program (e.g., different presets or different loudspeaker playback layouts). Since these configurations are static, one embodiment envisages measuring their loudness in the encoder (or before the encoding process) and introducing corresponding metadata fields, for example, in the MPEG-H 3DA stream.
[0092] However, as mentioned above, an important feature of modern audio coding systems such as MPEG-H 3DA is the support of user interactivity at the decoder. The user can, for example, adjust the volume of specific groups or switch them on or off. An important use case is dialog enhancement, where the user can manipulate the level of a dialog object or the groups associated with that dialog. In another embodiment, the user increases the level of an immersive acoustic bed represented by HOA-based groups. In another embodiment, the user may wish to switch to a specific group representing, for example, a video description for a hearing impaired track or a voice-over track.
[0093] Changing the level of a group also means that the overall loudness of the rendered audio scene is changed compared to the unmodified case. Therefore, a consistent playback loudness is no longer guaranteed after gain bidirectionality. As the user may change the levels of different objects more frequently, the loudness level of the audio output may change over time, even within the same program.
[0094] It is highly desirable to provide loudness control not only for the static representation of an audio program, but also to take into account user interactions that modify the loudness of an audio scene. The present invention allows for improved loudness control at the decoder to allow consistent loudness normalization even in the case of user interactions on the levels of groups of audio elements.
[0095] When a user changes the level of an audio element or group within a rendered audio scene, the program or preset loudness is preserved. In one embodiment, a loudness compensation gain is determined based on a reference loudness corresponding to the original audio scene and a modified loudness that takes into account the user's gain interaction. The loudness compensation gain is then applied to the rendered audio signal together with a standard loudness normalization gain to achieve the desired decoder target loudness.
[0096] 2 shows a schematic representation of an embodiment of an audio processor 1 - also called a decoder or simply a device for processing audio signals - which receives an audio signal 100 and provides an output signal 101. In this embodiment, the output signal 101 is an audio signal suitable for being fed to an amplifier (not shown) connected to a loudspeaker in a playback situation, or fed directly to a loudspeaker or headphones. The audio signal 100 comprises a bitstream having audio signals for individual audio objects and metadata providing information about the audio elements and how to handle them.
[0097] An audio signal 100 is provided to an audio signal modifier 2 which receives a user input 200. The user input 200 refers—in this embodiment—to the selection of at least one preset. A preset refers to a unique combination of a group of audio elements and an associated reference gain gi and / or group loudness Li of the corresponding group of audio elements. If the user does not select a preset, a default preset with default values will be used in this embodiment.
[0098] Furthermore, the user sets the gain values of the individual groups via the user input 200. A modified gain value hi means that the corresponding group will be amplified or attenuated corresponding to the reference gain value gi contained in the metadata. For example, a user may like to hear an amplified background choir but (unusually) not want to hear the leading voice. Therefore, the user may increase the gain value of the background choir and decrease the gain value of the leading voice, or switch this voice off.
[0099] The user also has the possibility to switch a group off or on. If the user does not want to hear a group, the group can be switched off. Conversely, if the metadata contains a flag that means that a group is switched off for a particular preset, the user can switch the group on. This may be the case, for example, if the audio signal contains different language versions of the dictated text and the presets point to different languages. Therefore, switching a group on or off indicates whether the group will be used in playback or not.
[0100] In short, the signal modification unit 2 modifies the audio signal 100 according to the user input 200 by amplifying or attenuating groups of audio elements belonging to the audio signal 100, and according to a selected preset or default preset covered by individual data of the metadata.
[0101] The signal modification unit 2 is followed by a configuration transformation unit 3 which transforms the data into the current playback configuration with which the audio signal 100 will be played. Which playback configuration is given and therefore the current situation is also covered by user input 200, e.g., selection from a list. For example, the metadata may refer to a surround sound situation, while the current playback state may allow stereo playback. This transformation, in one embodiment, refers to gain values as well as loudness values.
[0102] The configuration transformation unit 3 provides the transformed data to a loudness control unit 6 which also receives user input 200. Based on these data, the loudness control unit 6 calculates a loudness compensation gain C which is provided to the loudness manipulation unit 5.
[0103] The loudness manipulation unit 5 uses the loudness compensation gain C and the signal received from the mixer 4 to set the overall loudness of the output signal 101. The mixer 4 in this embodiment receives the audio signal 100 after modification by the audio signal modification unit 2 and transformation by the composition transformation unit 3 via the composition transformation unit 3 and combines different groups of audio elements (compare Fig. 1).
[0104] For purposes of explanation, the illustrated example considers the case where a unique audio scene is defined by a unique combination of presets, i.e., groups. Each group has an associated initial / default gain for a given preset. Furthermore, it is assumed that the loudness of each group within that preset is available. That preset may be selected by the user or may be set as the default preset. The following notations are used: Li is the loudness of the ith group of presets gi is the initial / default gain of the ith group (e.g. given in dB scale) hi is the corrected interactive (two-way) gain of the ith group (e.g., given in dB scale) M ref indicates a set of indices representing the groups included for the calculation of the reference loudness of a certain preset (or default audio scene). M mod denotes a set of indices representing the groups included for the calculation of the modified loudness of a certain preset (or modified audio scene).
[0105] If a group consists of a set of channel signals corresponding to a unique loudspeaker configuration, e.g., an HOA audio scene, multiple group loudness values may be included in the metadata. These different loudness values are associated with different loudspeaker configurations used for playback. For example, if a group represents a channel bed with a 5.1 or 22.2 loudspeaker configuration, a different loudness may be measured to play the group for the original 5.1 or 22.2 loudspeaker configuration compared to when the channel bed is to be mapped to a stereo playback system using a format converter. In this case, in one embodiment, the group loudness associated with stereo playback is selected if available in the transmitted metadata. Otherwise, the group loudness associated with the original loudspeaker configuration is used. A similar scheme is proposed for selecting an appropriate group loudness if a group represents an HOA-based audio scene. In this case, the group loudness associated with the current playback loudspeaker configuration should be used (if available in the metadata) instead of the group loudness associated with the reference loudspeaker layout.
[0106] In some embodiments, loudness information is not provided separately for each group, but rather the same loudness value is cited by a collection of groups.
[0107] In general, it is reasonable to assume that the audio signals of different groups are uncorrelated. In that case, the preset reference loudness can be calculated as follows: TIFF2026016762000002.tif21167
[0108] Similarly, the loudness of the modified audio scene is calculated as follows: TIFF2026016762000003.tif21167
[0109] If a group is switched off in the preset default settings, the group will have a reference loudness L ref Similarly, if the user switches off a group, that group is excluded when calculating the corrected loudness L mod If a group is switched off in the default preset but is switched on by the user in the modified scene, the corresponding group loudness Li is calculated based on the reference loudness L ref is excluded from the calculation of the corrected loudness L mod The exclusion of a switched-off group reduces its gain (gi or hi) by -∞ Note that this is interpreted as being equivalent to setting M ref =M mod Therefore, the loudness of both ref and L mod is calculated by referencing the same set of groups.
[0110] The loudness compensation gain C is the preset reference loudness L ref Preset corrected loudness L mod It is obtained by relating TIFF2026016762000004.tif27167
[0111] The loudness compensation gain C is, in one embodiment, limited within a certain range of allowed gain to avoid undesirable behavior for extreme cases. TIFF2026016762000005.tif27167
[0112] The loudness normalization gain G used for loudness normalization according to the prior art (see, for example, US Pat. No. 5,949,393) N is corrected according to the following formula: TIFF2026016762000006.tif16167This ensures consistent loudness after user gain interaction. Alternatively, loudness normalization can be applied to the original normalized gain G N and loudness compensation is performed based on a limited version of the compensation gain, C lim is performed separately on the audio signal using
[0113] The above explanation was based on a preset for an audio program. There may not always be presets available for a program, and only a single global default scene may be defined. This case is handled in the same way as for the presets mentioned above, where the index M ref and M mod The sets refer to a group of default scenes and their modified versions, respectively.
[0114] It may be appropriate to intentionally exclude certain groups from the loudness compensation process. For example, a certain group may be active for only a very short period of time within the program and completely silent for the rest of the time. Such a group may still have significant measurable loudness due to gating during the loudness measurement period, e.g., according to ITU-R BS.1770-3 - by the ITU Radiocommunication Sector (ITU-R) as one of the three sectors of the International Telecommunication Union (ITU). In that case, even though this group is only active for a very short period of time, this group loudness will affect the loudness compensation gain throughout the entire program duration. On the other hand, such sparse group signals may have only a negligible contribution to the loudness measurement of the entire program / preset mix.
[0115] For example, if the user chooses to boost such a sparse group / object, the loudness compensation will attenuate all remaining object elements for the entire program duration. Such behavior is undesirable and the loudness compensation process should ignore such special sparse groups. Therefore, the metadata contains a corresponding flag for this group to be ignored for the loudness compensation calculation.
[0116] To provide the above functionality, information is added to the metadata contained in the audio stream or audio signal that indicates whether a group should be excluded from loudness compensation, i.e., from the calculation of the reference loudness and modified loudness of a preset or global audio scene. In one embodiment, this information is a simple flag for each group that indicates whether the group is included in the loudness compensation process or not.
[0117] Different broadcasting regulations regarding loudness control use different approaches to defining program loudness: EBU-R128 requires measurement of the loudness of the full program mix, while ATSC A / 85 recommends measuring the loudness of only the anchor element of the program, which is typically represented by dialogue.
[0118] Such different approaches to measuring loudness for a program are also considered for loudness compensation. Anchor-based loudness compensation follows immediately from full-mix loudness compensation as described above.
[0119] For the anchor-based reference and modified loudness of a preset (or program default mix), only groups that contribute to the program anchor are included. The information about which groups are part of the program anchor is included in one embodiment in the metadata of the audio stream / audio signal. The reference loudness is given by: TIFF2026016762000007.tif21167 Here, A ref represents a set of indexes that indicate groups that are part of the anchor elements of the default audio scene or preset.
[0120] Similarly, the corrected loudness for anchor - based loudness compensation using the set A mod of group indexes (referring to groups that are part of the anchor elements of the modified audio scene or preset) is as follows. TIFF2026016762000008.tif21167
[0121] From the above, the compensation gain is obtained as follows. TIFF2026016762000009.tif27168
[0122] The remaining steps for performing loudness compensation are unchanged compared to the case of full - program mixing (as described above).
[0123] In some cases, a mixture of both loudness compensation approaches - anchor - based and full - program mixing - based - is beneficial for the user experience of loudness compensation.
[0124] In one embodiment, the anchor - based approach is used when one or all of the anchor groups are amplified by the user, i.e., when hi > gi. On the other hand, if the anchor group is attenuated, i.e., when hi < gi, then the loudness compensation for the full - mix loudness is used. Information about the anchor group is included in the metadata.
[0125] The loudness compensation approach described above requires information about the loudness of each group in a preset or global audio scene. In some scenarios, loudness information may be available only for some groups and missing for others. Thus, in one embodiment, the missing group loudness information is calculated from the loudness of the preset (or default audio scene) and the available group loudness values.
[0126] L p Let denote the measured loudness of the considered preset of the audio program, i.e. the measured joint loudness of the audio objects belonging to the individual presets. Furthermore, let B denote the set of indices to groups for which loudness information is available. The residual loudness L of a preset res is calculated from the preset loudness and available group loudness information and the default / initial gains of these groups. TIFF2026016762000010.tif21168
[0127] An alternative representation of the residual loudness may be obtained by considering unavailable group loudness values and corresponding default / initial gains. TIFF2026016762000011.tif21168
[0128] In practice, it is reasonable to assume that the loudness of each group for which loudness information is missing is equal. TIFF2026016762000012.tif16168
[0129] In this case, the residual loudness can be expressed as: TIFF2026016762000013.tif21168
[0130] From this, an estimate of the missing group loudness value is readily obtained from: TIFF2026016762000014.tif21168
[0131] Then, the reference loudness and modified loudness required for loudness compensation are calculated as already described above, and any missing group loudness Li is added to the corresponding estimated L A is replaced by
[0132] The estimation of the missing group loudness information is performed either at the encoder side or at the decoder side of the audio coding system.
[0133] If the estimation is performed on the encoder side, the information about group loudness in the metadata transmitted in the audio stream can be measured or alternatively a corresponding estimation as described above can be included, in which case the loudness compensation stage on the decoder side has all the necessary loudness information and can perform its processing according to the case where all group loudnesses have been pre-measured by the encoder.
[0134] If the estimation is performed in the decoder, the missing group loudness values in the metadata of the audio stream are estimated as described above, and loudness compensation is then performed based on the estimated group loudness values.
[0135] There are special use cases where no information about the loudness of any group is provided in the metadata of the audio stream. In this case, loudness compensation must operate only based on the available relevant rendering information, i.e., the default or initial gain gi of a group and its modified version hi after user interaction. This operation is called blind loudness compensation, since the loudness information for the groups is not known at the decoder side. In other embodiments, blind loudness compensation is also performed if only one group loudness is missing in the metadata.
[0136] For the compensation, the assumption is used that the loudness values of all groups in a preset are the same. In one embodiment of blind loudness compensation, M ref and M mod For all groups included in each, Li = L A This leads to the rule for calculating the loudness compensation gain according to the following formula: TIFF2026016762000015.tif27168
[0137] It should be noted that the gain factors for blind loudness compensation only require information about the group gains, but no loudness related information.
[0138] In a further embodiment, blind loudness compensation is performed if at least one group loudness is missing, thus blind loudness compensation is performed even if only one group loudness is missing.
[0139] This section summarizes what has been discussed above.
[0140] In one embodiment, a general set of indices is specified that indicates which groups should be included for the calculation of the reference loudness of a preset or default audio scene. This set is derived from information in the metadata of the audio stream and indicates whether a group should be included to perform loudness compensation for the default audio scene or preset. This information is typically introduced in the metadata of the audio stream at the encoder.
[0141] In the encoder, the loudness compensation process is controlled by appropriately defining these bitstream elements. For example, when a group should be excluded, the corresponding bitstream element is set to "false". In one embodiment, anchor-based loudness compensation is achieved by including only groups that are part of the anchor elements of the default audio scene or default preset, and setting the corresponding bitstream element to "true". Other ways of providing this information may be used in different configurations.
[0142] As already explained in one embodiment, if a group is switched off in the default audio scene or preset, the group will have a reference loudness L ref The resulting set of indices is K ref is shown as:
[0143] Similarly, any groups switched off in the modified scene will have a modified loudness L mod If a group is switched off in the default scene and switched on by the user in a modified scene, the corresponding group loudness is calculated by subtracting the reference loudness L ref is excluded from the calculation of the corrected loudness L mod is included in the calculation of the corrected loudness L mod The set of group indices of K mod It is shown as follows.
[0144] Then the loudness compensation gain is M ref K ref Replace with M mod K mod is replaced by , and the calculation is performed in the same manner as above.
[0145] If any of the group loudness information required to calculate either the reference or modified loudness is missing in the decoder, blind loudness compensation is used as a fallback mode. ref and K. mod ) applies in fallback mode.
[0146] 3 shows an embodiment of an audio encoder 20 that generates a single digital audio signal 100 based on different audio sources. The audio signal 100 includes metadata that may be used by, for example, the audio processors mentioned above.
[0147] The audio encoder 20 includes a loudness determiner 21 that determines a loudness value for at least one group having one or more audio elements 50. In the illustrated example, there are three audio sources X1, X2, and X3, all of which are included in one group. The loudness values of two of them, namely X2 and X3, are determined as L2 and L3 and supplied to a metadata writer 22. The metadata writer 22 introduces the determined loudness values for the two groups X2 and X3 into the metadata of the audio signal 100 as corresponding group reference loudness information L2 and L3.
[0148] Gain values as reference gains g1, g2, g3 for groups X1, X2 and X3 are also written by the metadata writer 22 into the metadata of the audio signal 100. According to further embodiments, group loudness and reference gain values are determined for specific presets and / or different playback configurations. Also, the individual global loudness L p The loudness for different presets as is also measured.
[0149] The loudness of the first audio element 50, marked as X1, is not measured by the loudness determination unit 21 but is calculated or estimated by the estimation unit 24 (see above) and supplied as a corresponding reference loudness L1 to the metadata writing unit 22 and written into the metadata.
[0150] In the illustrated embodiment, the controller 23 is connected to the loudness determiner 21 and the metadata writer 22. The controller 23 decides which groups should be taken into account or ignored for determining the loudness compensation gain C. An indication is written into the metadata by the metadata writer 22 as data regarding the usage of the groups. Corresponding data, for example in the form of flags, indicates which groups should be used or ignored for determining the loudness compensation gain C by the audio program or the decoder.
[0151] The resulting audio signal 100 contains the real signals received from the audio objects 50 and metadata characterizing these real signals and their intended treatment by the audio decoder 1. The metadata data refer to groups of audio objects, while it is also possible that one group covers only one audio object / element.
[0152] The metadata includes at least some of the following data: Measured loudness values for individual groups L i Reference gain value g for each group i a reference gain value g representing the loudness or prominence of each group relative to the union of other related groups; i Reference loudness L as the resulting loudness of the combined group for a given preset and / or a given playback configuration ref An indicator of whether or how a group or its corresponding values are used for determining the loudness compensation gain C (e.g., whether the group belongs to an anchor or whether the duration of the group is so short that it can be ignored) is used (e.g., for calculating the reference and / or modified loudness).
[0153] For each group, the metadata preferably includes different sets of data for different presets and / or different playback configurations, so that different recording and playback situations will result in different data sets for the associated group.
[0154] The present invention is described below through various embodiments that implement loudness compensation for user interaction using an audio coding system. On the encoder side, the loudness of each group included in the default audio scene and / or preset is determined. The loudness information is introduced into the audio stream or metadata included as part of the audio signal. Multiple loudness values are included for at least one group, with different values associated with different loudspeaker playback configurations (e.g., stereo, 5.1, or other). On the encoder side, additional metadata is generated corresponding to information whether a group should be included to perform loudness compensation, i.e. whether the group should be considered for the calculation of the reference loudness and the modified loudness, respectively. For example, anchor-based loudness compensation is realized by configuring the metadata to include only groups that are part of the anchor elements of the default audio scene or a given preset. A decoder receives an audio stream representing an audio signal and associated metadata, and decodes the audio stream to generate a decoded audio signal corresponding to a channel and / or object and / or higher-order Ambisonics format. Based on the metadata, the decoder selects all group indices to be included for loudness compensation of a given audio scene or preset. The decoder uses the reference loudness L of the audio scene or preset. ref is the default gain g for each selected group. i and the corresponding loudness information. If multiple loudness values are transmitted for a group, the loudness value associated with the given playback loudspeaker configuration is selected. Similarly, corrected loudness L mod is the loudness information of the selected group and the corrected gain h after user interaction. i It is calculated from The loudness compensation gain C for the default audio scene or preset is equal to the reference loudness L ref and corrected loudness L mod It is calculated based on: The loudness compensation gain C is applied to the audio signal before playback to provide the output signal.
[0155] In some embodiments, it is not possible for the encoder to measure the necessary loudness information for all groups. Therefore, the encoder calculates estimates of the missing group loudness values. The encoder may also apply different methods to estimate the missing (unmeasured) group loudness information. In that case, loudness compensation at the decoder is performed in the same way as if loudness information had been measured for all groups.
[0156] In a further embodiment, the audio stream has loudness information for only a limited number of groups, in which case the missing group loudness information is estimated at the decoder, in which case loudness compensation at the decoder side is performed in the same way as if all the necessary loudness information was included in the metadata of the audio stream.
[0157] Other embodiments include blind loudness compensation as a fallback mode in case the decoder lacks any required group loudness information to perform accurate loudness compensation. As mentioned above, a set of indices K is used to select the groups to be included in the reference and modified loudness calculations. ref and K. mod The same mechanism for determining K is used in the fallback mode. In other words, the set of group indices K ref and K. mod The selection of is still based on the corresponding information generated on the encoder side, which is provided together with the metadata of the audio stream.
[0158] Several embodiments of the present invention are described below, which can be combined with the above-mentioned embodiments.
[0159] A first embodiment refers to an audio processor for processing an audio signal, the audio processor comprising: an audio signal modification unit configured to modify the audio signal in response to a user input; a loudness control unit configured to determine a loudness compensation gain based on a reference loudness or a reference gain and based on a modified loudness or a modified gain, wherein the modified loudness or the modified gain depends on the user input; and a loudness manipulation unit configured to manipulate the loudness of a signal using the loudness compensation gain.
[0160] A second embodiment, which is dependent on the first embodiment, refers to an apparatus in which an audio signal comprises a bitstream with metadata, the metadata comprising a group loudness for a group and a gain value for a group.
[0161] A third embodiment, which is dependent on the first or second embodiment, refers to an apparatus, in which a loudness control unit is configured to calculate a reference loudness for a group or a set of groups using the group loudness or a plurality of group loudnesses and a reference gain value for the group or a reference gain value for the set of groups, and to calculate a modified loudness for a group or a set of groups using the group loudness or a plurality of group loudnesses and a modified gain value for the group or a modified gain value for the set of groups, wherein the modified gain value or a plurality of modified gain values is modified by user input.
[0162] A fourth embodiment, which depends on one of the preceding embodiments, refers to an apparatus, in which the loudness control unit is configured to exclude a group for determining the reference loudness if the group is switched off in the metadata of the audio signal, or to exclude a group for determining the modified loudness if the group is switched off in response to user input, or to exclude a group from the calculation of the reference loudness if the group is switched off in the metadata and switched on by user input, and vice versa.
[0163] A fifth embodiment, which depends on one of the preceding embodiments, refers to an apparatus, in which a loudness control unit is configured to calculate a loudness compensation gain by relating a reference loudness to a loudness of a preset, the preset including one or more groups, one group including one or more objects.
[0164] A sixth embodiment, which depends on any of the preceding embodiments, refers to an apparatus, in which a loudness control unit is configured to perform a limiting operation on the loudness compensation gain so that the loudness compensation gain is lower than an upper threshold or so that the loudness compensation gain is higher than a lower threshold.
[0165] A seventh embodiment, which relies on one of the preceding embodiments, refers to an apparatus, in which a loudness manipulation unit is configured to apply a gain to the signal determined by a loudness compensation gain and an original normalization gain determined by a target level set by an audio processor and a metadata level indicated in the metadata of the audio signal.
[0166] An eighth embodiment, which relies on one of the preceding embodiments, refers to an apparatus, in which an audio signal includes compensation metadata information indicating which groups should or should not be used for determining a loudness compensation gain, and a loudness control unit configured to use only the groups indicated by the compensation metadata information to be used for determining the loudness compensation gain, or to not use the groups indicated by the compensation metadata information not to be used for determining the loudness compensation gain.
[0167] A ninth embodiment, which relies on one of the preceding embodiments, refers to an apparatus, in which an audio signal is indicated to have anchor elements, and the loudness control unit is configured to use information about only one audio object or group of audio objects of the anchor element to determine a loudness compensation gain.
[0168] A tenth embodiment, which is dependent on one of the first to eighth embodiments, refers to an apparatus, in which an audio signal is indicated to have an anchor element, and the loudness control unit is configured to use only information about one audio object or a group of audio objects of the anchor element to determine a loudness compensation gain if one or more audio objects of the anchor element are amplified by user input, and to use information from one or more audio objects of the anchor element and information of one or more audio objects not included in the anchor element if one or more audio objects of the anchor element are attenuated by user input.
[0169] An eleventh embodiment, which depends on one of the preceding embodiments, refers to an apparatus, in which a loudness control unit is configured to calculate missing group loudness in an audio signal using the loudness of a preset including at least two groups and non-missing gain and loudness information for that preset.
[0170] A twelfth embodiment, which depends on one of the preceding embodiments, refers to an apparatus, in which a loudness control unit is configured to perform blind loudness compensation using one or more gain values for one or more groups and one or more modified gain values for one or more groups.
[0171] A thirteenth embodiment, which depends on one of the preceding embodiments, refers to an apparatus, in which a loudness control unit checks whether an audio signal includes reference loudness information, and if the audio signal does not include reference loudness information, performs blind loudness compensation using one or more reference gain values for one or more groups and one or more modified gain values for the one or more groups, or checks whether modified loudness information cannot be calculated, and if modified loudness information cannot be calculated, performs blind loudness compensation, where the blind loudness compensation includes using one or more reference gain values for one or more groups and one or more modified gain values for the one or more groups.
[0172] A fourteenth embodiment, which relies on one of the preceding embodiments, refers to an apparatus, in which an audio signal has different reference loudness information values for different playback configurations, the apparatus further comprising a format conversion unit that converts the signal into a predetermined playback configuration, and the loudness control unit is configured to select a specific loudness value for the specific playback configuration used by the format conversion unit.
[0173] The fifteenth embodiment refers to an audio encoder that generates an audio signal including metadata, and the audio encoder includes a loudness determination unit that determines loudness for a group having one or more audio objects, and a metadata writing unit that introduces the loudness for the group into the metadata as reference loudness information.
[0174] A sixteenth embodiment, which is dependent on the fifteenth embodiment, refers to an audio encoder, in which the loudness determination unit is configured to determine different loudness values for different playback configurations, and the metadata writing unit is configured to introduce the different loudness values into metadata for the different playback configurations.
[0175] A 17th embodiment, which is dependent on the 15th or 16th embodiment, refers to an audio encoder, wherein the audio encoder further comprises a controller that determines which groups should or should not be used for loudness compensation, and a metadata writing unit configured to write an indication into the metadata indicating which groups should or should not be used for loudness compensation.
[0176] An 18th embodiment, which depends on one of the 15th to 17th embodiments, refers to an audio encoder, in which the loudness determination unit is configured to calculate a group loudness value for a group, the group loudness value for which is missing in the metadata, and the metadata writing unit is configured to introduce the missing loudness value into the metadata, so that all groups of the audio signal have associated reference loudness information.
[0177] A 19th embodiment refers to a method for processing an audio signal, the method comprising the steps of modifying the audio signal in response to a user input, determining a loudness-compensation gain based on a reference loudness or a reference gain and based on a modified loudness or a modified gain, wherein the modified loudness or the modified gain is dependent on the user input, and manipulating the loudness of a signal using the loudness-compensation gain.
[0178] A twentieth embodiment refers to a method for generating an audio signal including metadata, comprising the steps of determining loudness for a group having one or more audio objects, and introducing the loudness for the group as reference loudness information into the metadata.
[0179] A twenty-first embodiment refers to a computer program for carrying out the method according to the nineteenth or twentieth embodiment when running on a computer or processor.
[0180] Although some aspects have been described above in the context of an apparatus, these aspects also represent a description of a corresponding method, and it is clear that a block or apparatus corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a corresponding block, item, or feature of a corresponding apparatus. Some or all of the method steps can be performed by (using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0181] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium, such as a wireless transmission medium like the Internet or a wired transmission medium.
[0182] Depending on certain implementation requirements, embodiments of the present invention can be configured in hardware or software. This configuration can be implemented using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, flash memory, etc., having electronically readable control signals stored therein that cooperate (or are capable of cooperating) with a programmable computer system to perform the methods of the present invention. Thus, the digital storage medium can be computer-readable.
[0183] Some embodiments according to the present invention include a data carrier having electronically readable control signals which are operable with a computer system programmable to carry out one of the methods described above.
[0184] Generally, embodiments of the present invention may be configured as a computer program product having program code operable to perform one of the methods of the present invention when the computer program product runs on a computer, the program code may for example be stored on a machine readable carrier.
[0185] Other embodiments of the invention comprise the computer program stored on a machine readable carrier for performing one of the methods described above.
[0186] In other words, an embodiment of the inventive method is therefore a computer program having a program code for performing one of the above described methods when the computer program runs on a computer.
[0187] Another embodiment of the invention is a data carrier (or digital storage medium, or computer readable medium) comprising a computer program recorded thereon for performing one of the methods described above. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0188] Another embodiment of the invention is, a data stream or a sequence of signals representing a computer program for performing one of the methods described above, which data stream or sequence of signals may be adapted to be transmitted via a data communication connection, for example the Internet.
[0189] Another embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described above.
[0190] Another embodiment comprises a computer having installed thereon the computer program for performing one of the methods described above.
[0191] Further embodiments of the present invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server that transmits the computer program to the receiver.
[0192] In some embodiments, a programmable logic device (such as a field programmable gate array) may be used to perform some or all of the functions of the methods described above. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described above. In general, such methods may be suitably performed by any hardware apparatus.
[0193] The above-described embodiments are merely illustrative of the principles of the present invention. It will be understood that modifications and variations of the above-described arrangements and details will be apparent to those skilled in the art. It is therefore the intention to be limited only by the subject matter of the claims appended hereto, and not by the specific details expressed in the manner of description and illustration of the embodiments. [remarks] [Claim 1] An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L moda loudness control unit (6) where the loudness compensation gain (C) or the modified gain (hi) depends on the user input, and the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) indicating which groups should or should not be used to determine the loudness compensation gain (C), the groups including one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising: [Claim 2] An audio processor (1) according to claim 1, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on at least one flag included in the metadata data; The flag indicates whether or how a group should be considered for determining the loudness compensation gain (C). [Claim 3] An audio processor (1) according to claim 1 or 2, the loudness control unit (6) is configured to use only the group for determining the loudness compensation gain (C) if the group belongs to an anchor included in the metadata of the audio signal (100). [Claim 4] An audio processor (1) according to claim 3, the loudness control unit (6) is configured to use only the group belonging to the anchor to determine the loudness compensation gain (C) if the modified gain (hi) of at least one group belonging to the anchor is greater than the corresponding reference gain (gi); and / or the loudness control unit (6) is configured to use the group belonging to the anchor and the group not belonging to the anchor to determine the loudness compensation gain (C) when the modified gain (hi) of at least one group belonging to the anchor is smaller than the corresponding reference gain (gi) and the modified gain (hi) depends on the user input. Audio processor. [Claim 5] An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod ) or the modified gain (hi) is dependent on the user input, and the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) that refers to at least one preset, the preset referring to a set of at least one group containing one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising: [Claim 6] An audio processor (1) according to claim 5, 5. The audio processor (1) configured according to any one of claims 1 to 4. [Claim 7] An audio processor (1) according to any one of claims 1 to 6, The loudness control unit (6) is configured to determine the loudness compensation gain (C) based on a group loudness (Li) and / or a gain value (gi) of the at least one group of the set referred to by the preset. [Claim 8] An audio processor (1) according to any one of claims 1 to 7, The loudness control unit (6) calculates the reference loudness (L) for the set mentioned by the preset using the individual group loudness (L i ) and the individual gain values (g i ). ref ), The loudness control unit (6) uses the individual group loudness (Li) and the individual modified gain value (hi) to calculate the modified loudness (L mod ), and the modified gain value (hi) is modified by the user input; Audio processor. [Claim 9] An audio processor (1) according to any one of claims 5 to 8, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata referring to a selected preset, The preset is selected by the user input. [Claim 10] An audio processor (1) according to any one of claims 5 to 9, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata referring to a default preset, An audio processor wherein the default preset is set prior to or independent of the user input. [Claim 11] An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) where the loudness compensation gain (C) or the modified gain (hi) is dependent on the user input, the loudness control unit (6) being configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) indicating whether a group is switched off or switched on, the group including one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising: [Claim 12] An audio processor (1) according to claim 11, 11. The audio processor (1) configured according to any one of claims 1 to 10. [Claim 13] An audio processor (1) according to claim 11 or 12, The loudness control unit (6) adjusts the modified loudness (L mod ) to determine a group of audio signals. [Claim 14] An audio processor (1) according to any one of claims 11 to 13, The loudness control unit (6) controls the reference loudness (Lref ) and if a group is switched on by the user input, the modified loudness (L mod ) configured to encompass the group to determine and / or The loudness control unit (6) controls the reference loudness (L ref ) and when a group is switched off by the user input, mod excluding the group to determine Audio processor. [Claim 15] An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) lacking at least one group loudness among a group of metadata included in the audio signal, wherein the modified gain (hi) depends on the user input; a loudness manipulation unit (5) configured to manipulate the loudness of the signal (101) using the loudness compensation gain (C); An audio processor comprising: [Claim 16] An audio processor (1) according to claim 15, 15. The audio processor (1) configured according to any one of claims 1 to 14. [Claim 17] An audio processor (1) according to claim 15 or 16, The loudness control unit (6) calculates the missing group loudness (Lp) using the preset loudness (Lp), the reference gain (gi) of the group having the missing group loudness, and the group loudness (Li) and the reference gain (gi) for the group having the group loudness (Li). A ). [Claim 18] An audio processor (1) according to any one of claims 15 to 17, The loudness control unit (6) is configured to determine the loudness compensation gain (C) using only at least one reference gain (gi) and at least one modified gain (hi) for blind loudness compensation when the metadata of the audio signal (100) is missing at least one group loudness. [Claim 19] An audio processor (1) according to any one of claims 15 to 18, The loudness control unit (6) is configured to determine the loudness compensation gain (C) using only at least one reference gain (gi) and at least one modified gain (hi) for blind loudness compensation when metadata of the audio signal (100) is invalid for group loudness. [Claim 20] An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) referring to a playback configuration for playback of the audio signal (100), wherein the modified gain (hi) is dependent on the user input; a loudness manipulation unit (5) configured to manipulate the loudness of the signal (101) using the loudness compensation gain (C); An audio processor comprising: [Claim 21] An audio processor (1) according to claim 20, 20. An audio processor (1) configured according to any one of claims 1 to 19. [Claim 22] An audio processor (1) according to claim 20 or 21, The loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata, which refers to a playback configuration and includes an associated group loudness (Li) and / or reference gain value (gi). [Claim 23] An audio processor (1) according to any one of claims 1 to 22, the audio signal (100) includes a bitstream having the metadata; and An audio processor, wherein the metadata includes the reference gain (gi) for at least one group. [Claim 24] An audio processor (1) according to any one of claims 1 to 23, An audio processor, wherein the metadata of the audio signal (100) includes a group loudness (Li) for at least one group. [Claim 25] An audio processor (1) according to any one of claims 1 to 24, The loudness control unit (6) calculates the reference loudness (L) for at least one group using the group loudness (L i ) and the gain value (g i ) of the at least one group. ref ), The loudness control unit (6) calculates the modified loudness (L mod ), the modified gain value (hi) is modified by the user input; Audio processor. [Claim 26] An audio processor (1) according to any one of claims 1 to 25, The loudness control unit (6) calculates the reference loudness (L) for each of the plurality of groups using the group loudness (L i ) and gain value (g i ). ref ), The loudness control unit (6) calculates the modified loudness (L) for each of the plurality of groups using the group loudness (L i ) and the modified gain value (h i ). mod ), Audio processor. [Claim 27] An audio processor (1) according to any one of claims 1 to 26, The loudness control unit (6) controls the loudness compensation gain (C) to exceed an upper threshold (C max ) and / or the loudness compensation gain (C) is lower than a lower threshold (C min) . [Claim 28] An audio processor (1) according to any one of claims 1 to 27, The loudness operation unit (5) controls the loudness compensation gain (C) and the normalization gain (G N ) and the corrected gain (G corrected ) to the signal, and the normalized gain (G N ) is determined by a target loudness level set by user input and a metadata loudness level included in the metadata of the audio signal (100). [Claim 29] An audio encoder (20) for generating an audio signal (100) including metadata, a loudness determiner (21) for determining a loudness value for at least one group having one or more audio elements (50); a metadata writing unit (22) that introduces the determined loudness value into the metadata as a group loudness (Li); an audio encoder including: [Claim 30] 30. An audio encoder (20) according to claim 29, the loudness determiner (21) is configured to determine different loudness values and / or different gain values for different playback configurations; and The metadata writer (22) is configured to introduce the determined different loudness values and / or different gain values into the metadata in association with the respective playback configurations. [Claim 31] An audio encoder (20) according to claim 29 or 30, the loudness determiner (21) is configured to determine different loudness values and / or different gain values for different presets referring to a set of at least one group containing one or more audio elements; and The metadata writing unit (22) is configured to introduce the determined different loudness values and / or different gain values into the metadata in association with respective presets. [Claim 32] An audio encoder (20) according to any one of claims 29 to 31, comprising: Further comprising a controller (23); the controller (23) is configured to determine which groups should be used or ignored for determining the loudness compensation gain (C); and The metadata writing unit (22) is configured to write an instruction into the metadata indicating which groups should be used or ignored for determining the loudness compensation gain (C). [Claim 33] An audio encoder (20) according to any one of claims 29 to 32, comprising: further comprising an estimator (24); the estimator (24) is configured to calculate a group loudness value for a group; the group loudness value for the group is not determined by the loudness determination unit (21), The metadata writer (22) is configured to introduce calculated group loudness values into the metadata such that all groups in the audio signal (100) have an associated group loudness. [Claim 34] A method for processing an audio signal (100), comprising: modifying the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref) or based on the reference gain (gi), and on the other hand on the corrected loudness (L mod ) or the modified gain (hi), determining a loudness compensation gain (C), mod ) or the modified gain (hi) is dependent on the user input; the loudness compensation gain (C) is determined based on metadata of the audio signal (100) indicating whether a group contained in the audio signal (100) should or should not be used to determine the loudness compensation gain (C), the group comprising one or more audio elements; and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) that refers to a preset, the preset referring to a set of at least one group containing one or more audio elements; and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) indicating whether a group is switched off or switched on, the group comprising one or more audio elements; and / or The loudness compensation gain (C) is a value obtained by adjusting at least one group loudness (L) in one group of metadata included in the audio signal (100). A ) is determined based on metadata of the audio signal (100) in a missing state, and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) referring to a playback configuration for reproduction of the audio signal (100); manipulating the loudness of a signal using the loudness compensation gain (C); A method for providing [Claim 35] A method for generating an audio signal (100) including metadata, comprising: determining a loudness value for a group having one or more audio elements; introducing the loudness value determined for the group into the metadata as group loudness (Li); A method comprising: [Claim 36] 36. A computer program for carrying out the method of claim 34 or 35 when running on a computer or processor.
Claims
1. An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) where the loudness compensation gain (C) or the modified gain (hi) depends on the user input, and the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) indicating which groups should or should not be used to determine the loudness compensation gain (C), the groups including one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising:
2. An audio processor (1) according to claim 1, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on at least one flag included in the metadata data; The flag indicates whether or how a group should be considered for determining the loudness compensation gain (C).
3. An audio processor (1) according to claim 1 or 2, The loudness control unit (6) is configured to use only a group for determining the loudness compensation gain (C) if the group belongs to an anchor included in the metadata of the audio signal (100).
4. An audio processor (1) according to claim 3, the loudness control unit (6) is configured to use only the groups belonging to the anchor for determining the loudness compensation gain (C) if the modified gain (hi) of at least one group belonging to the anchor is greater than the corresponding reference gain (gi); and / or the loudness control unit (6) is configured to use a group belonging to the anchor and a group not belonging to the anchor to determine the loudness compensation gain (C) when a modified gain (hi) of at least one group belonging to the anchor is smaller than the corresponding reference gain (gi) and the modified gain (hi) depends on the user input. Audio processor.
5. An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod ) or the modified gain (hi) is dependent on the user input, and the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) that refers to at least one preset, the preset referring to a set of at least one group containing one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising:
6. An audio processor (1) according to claim 5, 5. The audio processor (1) configured according to any one of claims 1 to 4.
7. An audio processor (1) according to any one of claims 1 to 6, The loudness control unit (6) is configured to determine the loudness compensation gain (C) based on a group loudness (Li) and / or a gain value (gi) of the at least one group of the set referred to by the preset.
8. An audio processor (1) according to any one of claims 1 to 7, The loudness control unit (6) uses the individual group loudness (Li) and the individual gain values (gi) to calculate the reference loudness (L ref ) configured to determine The loudness control unit (6) uses the individual group loudness (Li) and the individual modified gain value (hi) to calculate the modified loudness (L mod ) and the modified gain value (hi) is modified by the user input; Audio processor.
9. An audio processor (1) according to any one of claims 5 to 8, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata referring to a selected preset, The audio processor, wherein the preset is selected by the user input.
10. An audio processor (1) according to any one of claims 5 to 9, the loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata referring to a default preset, An audio processor wherein the default preset is set prior to or independent of the user input.
11. An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) where the loudness compensation gain (C) or the modified gain (hi) is dependent on the user input, the loudness control unit (6) being configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) indicating whether a group is switched off or switched on, the group comprising one or more audio elements; a loudness manipulation unit (5) configured to manipulate the loudness of a signal using the loudness compensation gain (C); An audio processor comprising:
12. An audio processor (1) according to claim 11, The audio processor (1) is configured according to any one of claims 1 to 10.
13. An audio processor (1) according to claim 11 or 12, The loudness control unit (6) adjusts the modified loudness (L mod ) excluding the group.
14. An audio processor (1) according to any one of claims 11 to 13, The loudness control unit (6) adjusts the reference loudness (L ref ) and if a group is switched on by the user input, the modified loudness (L mod ) configured to encompass the group to determine and / or The loudness control unit (6) adjusts the reference loudness (L ref ) and when a group is switched off by the user input, mod and excluding the group to determine Audio processor.
15. An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) lacking at least one group loudness among a group of metadata included in the audio signal, wherein the modified gain (hi) depends on the user input; a loudness manipulation unit (5) configured to manipulate the loudness of the signal (101) using the loudness compensation gain (C); An audio processor comprising:
16. 16. An audio processor (1) according to claim 15, The audio processor (1) is configured according to any one of claims 1 to 14.
17. An audio processor (1) according to claim 15 or 16, The loudness control unit (6) calculates the missing group loudness (Lp) using a preset loudness (Lp), a reference gain (gi) of a group having a missing group loudness, and a group loudness (Li) and a reference gain (gi) for a group having a group loudness (Li). A ).
18. An audio processor (1) according to any one of claims 15 to 17, The loudness control unit (6) is configured to determine the loudness compensation gain (C) using only at least one reference gain (gi) and at least one modified gain (hi) for blind loudness compensation when the metadata of the audio signal (100) is missing at least one group loudness.
19. An audio processor (1) according to any one of claims 15 to 18, The loudness control unit (6) is configured to determine the loudness compensation gain (C) using only at least one reference gain (gi) and at least one modified gain (hi) for blind loudness compensation when metadata of the audio signal (100) is invalid for group loudness.
20. An audio processor (1) for processing an audio signal (100), comprising: an audio signal modifying unit (2) configured to modify the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod a loudness control unit (6) configured to determine a loudness compensation gain (C) based on the corrected loudness (L mod a loudness control unit (6) configured to determine the loudness compensation gain (C) based on metadata of the audio signal (100) referring to a playback configuration for reproducing the audio signal (100), wherein the modified gain (hi) is dependent on the user input; a loudness manipulation unit (5) configured to manipulate the loudness of the signal (101) using the loudness compensation gain (C); An audio processor comprising:
21. 21. An audio processor (1) according to claim 20, 20. An audio processor (1) configured according to any of claims 1 to 19.
22. 22. An audio processor (1) according to claim 20 or 21, The loudness control unit (6) is configured to determine the loudness compensation gain (C) based on data of the metadata that refers to a playback configuration and includes associated group loudness (Li) and / or reference gain values (gi).
23. An audio processor (1) according to any one of the preceding claims, comprising: the audio signal (100) comprises a bitstream having the metadata; and An audio processor, wherein the metadata includes the reference gain (gi) for at least one group.
24. An audio processor (1) according to any one of claims 1 to 23, An audio processor, wherein the metadata of the audio signal (100) includes a group loudness (Li) for at least one group.
25. An audio processor (1) according to any one of the preceding claims, comprising: The loudness control unit (6) calculates the reference loudness (L) for the at least one group using the group loudness (Li) and the gain value (gi) of the at least one group. ref ) configured to determine The loudness control unit (6) calculates the modified loudness (L mod ) configured to determine the modified gain value (hi) is modified by the user input; Audio processor.
26. An audio processor (1) according to any one of claims 1 to 25, The loudness control unit (6) calculates the reference loudness (L) for each of the plurality of groups using the group loudness (L i ) and gain value (g i ) of each of the plurality of groups. ref ) configured to determine The loudness control unit (6) calculates the modified loudness (L) for each of the groups using the group loudness (L i ) and the modified gain value (h i ). mod ) configured to determine Audio processor.
27. An audio processor (1) according to any one of the preceding claims, comprising: The loudness control unit (6) controls the loudness compensation gain (C) to be equal to or lower than an upper threshold (C max ) and / or the loudness compensation gain (C) is lower than a lower threshold (C min ) .
28. An audio processor (1) according to any one of the preceding claims, comprising: The loudness operation unit (5) controls the loudness compensation gain (C) and the normalization gain (G N ) and the corrected gain (G corrected ) to the signal, and the normalized gain (G N ) is determined by a target loudness level set by user input and a metadata loudness level included in the metadata of the audio signal (100).
29. An audio encoder (20) for generating an audio signal (100) including metadata, comprising: a loudness determiner (21) for determining a loudness value for at least one group having one or more audio elements (50); a metadata writing unit (22) that introduces the determined loudness value into the metadata as a group loudness (Li); an audio encoder including:
30. 30. An audio encoder (20) according to claim 29, the loudness determiner (21) is configured to determine different loudness values and / or different gain values for different playback configurations; and The metadata writer (22) is configured to introduce the determined different loudness values and / or different gain values into the metadata in association with the respective playback configurations.
31. 31. An audio encoder (20) according to claim 29 or 30, the loudness determiner (21) is configured to determine different loudness values and / or different gain values for different presets referring to a set of at least one group containing one or more audio elements; and The metadata writing unit (22) is configured to introduce the determined different loudness values and / or different gain values into the metadata in association with respective presets.
32. An audio encoder (20) according to any one of claims 29 to 31, further comprising a controller (23); the controller (23) is configured to determine which groups should be used or ignored for determining the loudness compensation gain (C); and The metadata writing unit (22) is configured to write an instruction into the metadata indicating which groups should be used or ignored for determining the loudness compensation gain (C).
33. An audio encoder (20) according to any one of claims 29 to 32, comprising: further comprising an estimator (24); the estimator (24) is configured to calculate a group loudness value for a group; the group loudness value for the group is not determined by the loudness determination unit (21), The metadata writer (22) is configured to introduce calculated group loudness values into the metadata such that all groups of the audio signal (100) have an associated group loudness.
34. A method for processing an audio signal (100), comprising: modifying the audio signal (100) in response to user input; On the other hand, the reference loudness (L ref ) or the reference gain (gi), and on the other hand the corrected loudness (L mod ) or the modified gain (hi), determining a loudness compensation gain (C), mod ) or the modified gain (hi) is dependent on the user input; the loudness compensation gain (C) is determined based on metadata of the audio signal (100) indicating whether a group contained in the audio signal (100) should or should not be used to determine the loudness compensation gain (C), the group comprising one or more audio elements; and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) that refers to a preset, the preset referring to a set of at least one group containing one or more audio elements; and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) indicating whether a group is switched off or switched on, the group comprising one or more audio elements; and / or The loudness compensation gain (C) is determined by adjusting at least one group loudness (L) among the metadata of one group included in the audio signal (100). A ) is determined based on metadata of the audio signal (100) in a missing state, and / or the loudness compensation gain (C) is determined based on metadata of the audio signal (100) referring to a playback configuration for reproduction of the audio signal (100); manipulating the loudness of a signal using said loudness compensation gain (C); A method for providing
35. A method for generating an audio signal (100) including metadata, comprising: determining a loudness value for a group having one or more audio elements; introducing the loudness value determined for the group into the metadata as group loudness (Li); A method comprising:
36. 36. A computer program for carrying out the method of claim 34 or 35 when running on a computer or processor.
Citation Information
Patent Citations
Decoder, encoder and method for informed loudness estimation in object-based audio coding systems
EP2879131A1