Method and apparatus for processing an initial audio signal

By optimizing the processing of audio signals with signal modifiers and perception models, the problem of maintaining sound aesthetics and improving speech clarity is solved, and cost-effective speech clarity is achieved to adapt to different hearing impairments and personal preferences.

CN115699172BActive Publication Date: 2025-07-08FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202080101547.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-29
Publication Date
2025-07-08
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

The prior art is difficult to improve the speech clarity of audio and audio-visual media while maintaining the acoustic aesthetic, especially for people with hearing impairment, and the existing methods are costly or ineffective.

Method used

By receiving the initial audio signal, modifying using the first and second signal modifiers, evaluating and selecting modified audio signals that satisfy perceived similarity and speech clarity, the processing process is optimized using neural networks and perceived models to ensure the optimal tradeoffs for speech clarity and sound scenes.

Benefits of technology

It realizes improving speech clarity without significantly changing the sound aesthetics, reducing processing costs and workload, adapting to different hearing impairments and personal preferences, and providing a more efficient method of improving speech clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699172B_ABST
    Figure CN115699172B_ABST
Patent Text Reader

Abstract

A method (100) for processing an initial audio signal (AS) including a target portion (AS_TP) and a side portion (AS_SP), comprising the steps of: receiving the initial audio signal (AS); modifying the received initial audio signal (AS) by using a first signal modifier to obtain a first modified (110a) audio signal, and modifying the received initial audio signal (AS) by using a second signal modifier to obtain a second modified audio signal (Second MOD AS); comparing the received initial audio signal (AS) with the first modified audio signal (First MOD AS) to obtain a first perceptual similarity value (First PSV), the First PSV describing the perceptual similarity between the initial audio signal (AS) and the first modified audio signal (First MOD AS); and comparing the received initial audio signal (AS) with the second modified audio signal (Second MOD AS) to obtain a second perceptual similarity value (Second PSV), the Second PSV describing the perceptual similarity between the initial audio signal (AS) and the second modified audio signal (Second MOD AS); and selecting (130) the first modified audio signal (First MOD AS) or the second modified audio signal (Second MOD AS) depending on the respective first perceptual similarity value or second perceptual similarity value (Second PSV).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a method and a corresponding apparatus for processing an initial audio signal (such as a recording or raw data). Preferred embodiments relate to a method (way and algorithm) for improving speech intelligibility and for listening to broadcast audio materials. Background Art

[0002] When producing and broadcasting audio and audiovisual media (such as movies, TV, radio, podcasts, YouTube videos), it is not always possible to ensure a sufficiently high speech intelligibility in the final mix, for example due to the addition of too much background sound (music, sound effects, noise in the recording, etc.).

[0003] This is particularly problematic for people with hearing impairments, but improving speech intelligibility is also beneficial for people with normal hearing or non-native listeners.

[0004] A fundamental problem in the production of audio and audiovisual media is that the background signals (music, sound effects, atmosphere) constitute an important part of the sound aesthetics in the production, that is, the background signals cannot be regarded as "interference noise" that should be eliminated as much as possible. Therefore, all methods aimed at improving speech intelligibility or reducing listening effort for this application should additionally consider changing the originally expected sound characteristics as little as possible to take into account the high-quality requirements and creative aspects of sound production. However, currently, there is no technical method or tool for ensuring the best compromise between good intelligibility and maintaining the sound scene / recording.

[0005] However, there are different technical methods that can basically improve the speech intelligibility (or reduce the listening effort) of audio and audiovisual media:

[0006] One solution is to have professional sound engineers manually produce alternative audio mixes so that end users can freely choose between the original mix and the mix with improved speech intelligibility. For example, by simulating hearing loss and ensuring that the expected mix is also suitable for listeners with target hearing loss, a mix with improved intelligibility can be produced [1]. However, this manual process is very costly and not applicable to most of the produced audio / visual media.

[0007] As an alternative solution for providing automatic signal enhancement, there are different methods for reducing or eliminating unwanted signal parts (such as interference noise), however, these methods are different from the technical method of the present invention:

[0008] Improving speech clarity by interference noise reduction methods for mixed signals: Such methods are aimed at processing a mixed signal that includes both a target signal (e.g., speech) and an interfering signal (e.g., background noise), such that as much interfering noise as possible is eliminated while the target signal ideally remains unchanged (e.g., according to the method of [2]). Since these methods must estimate the respective parts of the target and interfering noise components in the mixed signal, these methods are always based on assumptions about the physical properties of the signal components. Such algorithms are used, for example, in hearing aids and mobile phones, belong to the prior art and will continue to be further developed.

[0009] In the past few years, more and more machine learning (neural network)-based methods have been proposed that aim to separate different sources in a mixed signal. Based on large amounts of data, these methods are trained for specific problems (e.g., separating several speakers in a mix [3]), and can basically be used to extract dialogue from the ambiance / music in audiovisual media, thus providing a basis for remixing with improved SNR. In [4], such a method has been proposed for allowing the user to select and adjust the ratio of speech to background by themselves.

[0010] Improving speech clarity by preprocessing the speech signal: In some applications, the target signal (e.g., speech) is separated from other signal parts; thus, the target signal is not a mixed signal as described above, and the method does not require any estimation of the signal components corresponding to the target and interfering noise. For example, this is the case for train station announcements. At the same time, at the signal processing level, the interfering noise cannot be affected, i.e., the interfering noise cannot be eliminated or reduced (e.g., the noise of passing trains interferes with the clarity of station announcements). For such application scenarios, there are the following methods: preprocessing the target signal adaptively such that the clarity of the target signal is optimal or improved in the currently present interfering noise (e.g., the method of [5]). Such a method uses, for example, band-pass filtering of the target signal, frequency-dependent amplification, time delay, and / or dynamic compression, and will basically also be applicable to audiovisual media without (significantly) modifying the background noise / ambiance.

[0011] Encoding the target and background noise as separate audio objects: Furthermore, there are the following methods: when encoding and transmitting an audio signal, the information about the target signal is encoded parametrically such that the energy of the target signal can be adjusted separately during decoding at the receiver. Increasing the energy of the target object (e.g., speech) relative to other audio objects (e.g., ambiance) can lead to improved speech clarity

[11] .

[0012] Detection and level adaptation of the speech signal in a mixed signal: On top of this, there is the following technical system: identifying the speech channels in the mixed signal and modifying these channels with the aim of obtaining improved speech intelligibility, for example increasing the volume of these channels. Depending on the type of modification, this will only improve speech intelligibility if there is no other interfering noise present in the mixed signal at the same time

[12] .

[0013] Reducing channels that mainly do not include speech: In a multi-channel audio signal mixed in such a way that one channel (usually the center) includes most of the speech information while the other channels (e.g., left / right) mainly include background noise, a technical solution consists of attenuating the non-speech channels by a fixed gain (e.g., 6 dB) and in this way improving the signal-to-noise ratio (e.g., the adapted downmix rules of a sound retrieval system (SRS) dialogue clarity or a surround sound decoder).

[0014] In this method, it may happen that: parts of the background noise that are already very low and actually have no adverse effect on speech intelligibility are also attenuated. This may reduce the overall sound aesthetic impression as the atmosphere intended by the sound engineer can no longer be perceived. To prevent this, US 8,577,676 B2 describes a method in which the non-speech channels are only reduced to the effect that the measure of speech intelligibility reaches a specific threshold, but no more. Additionally, US 8,577,676 B2 discloses a method in which multiple frequency-dependent attenuations are calculated, each having the effect that the measure of speech intelligibility reaches a specific threshold. Then, the option that maximizes the loudness of the background noise is selected from among the multiple options. This is based on the assumption that this best preserves the original sound characteristics as much as possible.

[0015] Based on this, US 2016 / 0071527 A1 describes a method in which, contrary to the general assumption, when the non-speech channels also include relevant speech information and thus reducing them may be detrimental to clarity, the non-speech channels are not reduced or not reduced much. This document also includes a method in which multiple frequency-dependent attenuations are calculated and the attenuation that maximizes the loudness of the background noise is selected (again based on the assumption that this best preserves the original sound characteristics as much as possible).

[0016] Both US patent documents describe in their independent claims very specific methods that are not required for the invention described herein (e.g., scaling the reduction factor with the probability of speech occurrence). Therefore, the present invention can be implemented without using the techniques disclosed in US 8,577,676 B2 and US 2016 / 0071527 A1.

[0017] US 8,195,454 B2 describes a method of detecting portions of an audio signal where speech occurs by using voice activity detection (VAD). Then, one or more parameters are modified for these portions (e.g., dynamic range control, dynamic equalization, spectral sharpening, frequency shifting, voice extraction, noise reduction, or other speech enhancement actions) such that a measure of speech intelligibility (e.g., the speech intelligibility index (SII) [6]) is maximized or increased above a desired threshold. Here, hearing loss or the preferences of the listener or the noise in the listening environment can be considered.

[0018] US 8,271,276 B1 describes the loudness or level adaptation of speech segments, where the amplification factor depends on previous time segments. This is not relevant to the core of the invention described herein and would only become relevant if the invention described herein simply changes the loudness or level of segments identified as speech depending on previous segments. It does not include the adaptation of the audio signal other than amplifying the speech segments, such as source separation, reducing background noise, spectral changes, dynamic compression. Thus, the steps disclosed in US 8,271,276 B1 are not adverse either.

[0019] The object of the present invention is to provide a concept that achieves an improved compromise between (speech) intelligibility and maintaining the sound scene.

[0020] This object is achieved by the content of the independent claims.

[0021] An embodiment of the present invention provides a method for processing an initial audio signal including a target portion (e.g., a speech portion) and a side portion (e.g., ambient noise). The method includes the following four steps:

[0022] 1. Receiving the initial audio signal;

[0023] 2. Modifying the received initial audio signal by using a first signal modifier to obtain a first modified audio signal, and modifying the received initial audio signal by using a second signal modifier to obtain a second modified audio signal;

[0024] 3. Evaluating the first modified audio signal against an evaluation criterion to obtain a first evaluation value describing the degree of satisfaction of the evaluation criterion, and evaluating the second modified audio signal against the evaluation criterion to obtain a second evaluation value describing the degree of satisfaction of the evaluation criterion;

[0025] 4. Selecting the first modified audio signal or the second modified audio signal depending on the corresponding first evaluation value or second evaluation value.

[0026] According to an embodiment, the evaluation criteria can be one or more of a group including perceptual similarity, speech clarity, loudness, sound pattern, and spatial sense. Note that according to an embodiment, the selection step can be performed based on a plurality of independent first evaluation values and second evaluation values that describe independent evaluation criteria. The evaluation criteria and in particular the selection step can depend on a so-called optimization goal. Thus, according to an embodiment, the method includes the following steps: receiving information about an optimization goal that defines personal preferences; wherein the evaluation criteria depend on the optimization goal; or wherein the modification and / or evaluation and / or selection steps depend on the optimization goal; or wherein the weighting of the independent first evaluation values and second evaluation values of the independent evaluation criteria used for the selection step depends on the optimization goal.

[0027] For example, if the optimization goal is a combination of two elements (e.g., the best speech clarity and a tolerable perceptual similarity between an initial audio signal and a modified audio signal), then the weighting for the selection can be performed. For example, the two criteria of speech clarity and perceptual similarity can be evaluated separately such that the corresponding evaluation values of the evaluation criteria are determined, whereupon the selection is then performed based on the weighted evaluation values. The weighting depends on the optimization goal and vice versa, and can be set by personal preferences.

[0028] According to an embodiment, the adaptation, evaluation, and selection steps can be performed by using a neural network / artificial intelligence.

[0029] According to a preferred embodiment, it is assumed that the speech clarity is improved in a sufficient manner by two or more modifiers used. Expressed from another perspective, this means that only modifiers that can improve the speech clarity high enough or the clarity of the output speech is sufficient are considered. In the next step, a selection is made between the differently modified signals. For this selection, the perceptual similarity is used as an evaluation criterion, so that steps 3 and 4 (see the above method) can be performed as follows:

[0030] 3. Comparing the received initial audio signal with a first modified audio signal to obtain a first perceptual similarity value that describes the perceptual similarity between the initial audio signal and the first modified audio signal; and comparing the received initial audio signal with a second modified audio signal to obtain a second perceptual similarity value that describes the perceptual similarity between the initial audio signal and the second modified audio signal; and

[0031] 4. Selecting the first modified audio signal or the second modified audio signal depending on the corresponding first perceptual similarity value or second perceptual similarity value. Summary of the Invention

[0032] According to an embodiment of the present invention, when the first perceptual similarity value is higher than the second perceptual similarity value (a high first perceptual similarity value indicates a higher perceptual similarity of the first modified audio signal), the first modified audio signal is selected; vice versa, when the second perceptual similarity value is higher than the first perceptual similarity value (a high second perceptual similarity value indicates a higher perceptual similarity of the second modified audio signal), the second modified audio signal is selected. According to a further embodiment, instead of the perceptual similarity value, another value, such as a loudness value, may be used.

[0033] According to a further embodiment, this adaptation method with the comparison step 3 based on the perceptual similarity value and the selection step 4 can be enhanced by an additional step of evaluating the first modified signal and the second modified signal for another optimization criterion (e.g., for speech intelligibility) after step 2 and before step 3. As described above, in this case some of the modified signals may not be considered because, for example, the first evaluation criterion is not (sufficiently) met when the speech intelligibility is too low. Alternatively, all evaluation criteria may be considered during the selection step, either unweighted or weighted. This weighting can be selected by the user.

[0034] According to an embodiment, the method further comprises the step of outputting the first modified audio signal or the second modified audio signal depending on the selection.

[0035] An embodiment of the present invention provides a method, wherein the target part is the speech part of the initial audio signal, and the side part is the ambient noise part of the audio signal.

[0036] Embodiments of the present invention vary based on defining different speech intelligibility options regarding their improvement effects, depending on multiple influencing factors, e.g., depending on the input audio stream or the input audio scene. In one audio stream, the best speech intelligibility algorithm may also vary by scene. Therefore, embodiments of the present invention analyze different modifications of the audio signal, particularly regarding the perceptual similarity between the initial audio signal and the modified audio signal, in order to select the modifier / modified audio signal with the highest perceptual similarity. The system / concept makes it possible for the first time to perceptually change the overall sound only when necessary, but as little as possible, to meet two requirements, namely improving the speech intelligibility of the initial signal (or reducing the listening effort), while minimally affecting the sound aesthetic component. Compared to non-automatic methods, this represents a significant reduction in workload and cost, and compared to methods that have so far only been used as boundary conditions for improving intelligibility, this represents a significant added value. Since maintaining this sound aesthetic represents an important part of user acceptance, this has not been considered in automated methods so far.

[0037] According to an embodiment, when the corresponding first perceptual similarity value or second perceptual similarity value is below a threshold, the step of outputting the initial audio signal instead of outputting the first modified audio signal or the second modified audio signal is performed. "Below" indicates that the modified signal is not similar enough to the initial audio signal. This is advantageous because the system can automatically check the mix for speech intelligibility or listening effort, and at the same time it ensures that the overall sound is perceptually changed in an efficient manner.

[0038] Embodiments of the present invention provide a method, wherein the comparing step comprises: extracting the first perceptual similarity value and / or the second perceptual similarity value by using a (perceptual) model such as the PEAQ model, the POLQA model, and / or the PEMO-Q model [8], [9],

[10] . Note that PEAQ, POLQA, and PEMOQ are specific models that are trained to output the perceptual similarity of two audio signals. According to an embodiment, the degree of processing is controlled by another model.

[0039] Note that, according to an embodiment, the first perceptual similarity value and / or the second perceptual similarity value depends on the physical parameters of the first modified audio signal or the second modified audio signal, the volume level of the first modified audio signal or the second modified audio signal, the psychoacoustic parameters of the first modified audio signal or the second modified audio signal, the loudness information of the first modified audio signal or the second modified audio signal, the pitch information of the first modified audio signal or the second modified audio signal, and / or the perceptual source width information of the first modified audio signal or the second modified audio signal.

[0040] Embodiments of the present invention provide a method, wherein the first signal modifier and / or the second signal modifier is configured to perform an SNR increase (e.g., of the initial audio signal), dynamic compression (e.g., of the initial audio signal); and / or wherein, if the initial audio signal includes a separate target portion and a separate side portion, the modifying step comprises: increasing the target portion, increasing the frequency weighting of the target portion, dynamically compressing the target portion, reducing the side portion, reducing the frequency weighting of the side portion; alternatively, if the initial audio signal includes a combined target portion and side portion, the modifying comprises: performing a separation of the target portion and the side portion. Generally, this means that embodiments of the present invention provide a method, wherein the first modified audio signal and / or the second modified audio signal comprises: a target portion moved to the foreground and a side portion moved to the background, and / or a speech portion moved to the foreground as the target portion and an ambient noise portion moved to the background as the side portion.

[0041] According to an embodiment, the step of selection is performed taking into account one or more additional factors, such as the hearing loss level of the hearing-impaired person, the individual hearing performance; individual frequency-related hearing performance; individual preferences; individual preferences regarding the signal modification rate. Similarly, according to an embodiment, the step of modification and / or comparison is performed taking into account one or more factors, such as the hearing loss level of the hearing-impaired person, the individual hearing performance; individual frequency-related hearing performance; individual preferences; individual preferences regarding the signal modification rate. Thus, selection, modification, and / or comparison can also take into account individual hearing or individual preferences.

[0042] According to an embodiment, the model for controlling the processing can be configured, for example, for hearing loss or individual preferences.

[0043] According to an embodiment, the step of comparison is performed for: the entire initial audio signal and the entire first modified audio signal and second modified audio signal, or the target part of the individual audio signals is compared with the corresponding target parts of the first modified audio signal and the second modified audio signal, or the side part of the initial audio signal is compared with the side parts of the first modified audio part and the second modified audio part.

[0044] Embodiments of the present invention provide a method, wherein the method further includes the following initial steps: analyzing an initial audio part to determine a speech part; comparing the speech part with an ambient noise part to evaluate the speech intelligibility of the initial audio signal, and if the value indicating the speech intelligibility is lower than a threshold, activating a first signal modifier and / or a second signal modifier for the modification step. Thus, it is advantageous to perform processing only at the channels where speech appears. Here, a modified mix is generated for the speech part, wherein the mix is designed to meet or maximize a specific perceptual metric.

[0045] Embodiments of the present invention provide a method, wherein the initial audio signal includes a plurality of time frames or scenes, and wherein the basic steps are repeated for each time frame or scene.

[0046] According to an embodiment, a first modifier can be used to adapt the first time frame, and another modifier is selected for the second time frame. To ensure perceptual continuity, a transition can be inserted between the time frame or the adapted parts of two time frames. For example, the end of the first time frame and the start of the subsequent time frame are adapted for their adaptation performance. For example, an interpolation between two adaptation methods can be applied. According to a further embodiment, the same modifier can be used for all or multiple subsequent time frames in order to achieve perceptual continuity. According to a further embodiment, the adaptation of the time frame can be performed even if, for example, it is not required from the perspective of clarity performance. However, this can ensure the perceptual similarity between the corresponding time frames.

[0047] An embodiment of the present invention provides a computer program having program code for performing the above method when run on a computer.

[0048] Another embodiment of the present invention provides an apparatus for processing an initial audio signal. The apparatus includes: an interface for receiving the initial audio signal; a corresponding modifier for processing the initial audio signal to obtain a corresponding modified audio signal, an evaluator for performing an evaluation of the corresponding modified audio signal, and a selector for selecting a first modified audio signal or a second modified audio signal depending on a corresponding first evaluation value or a second evaluation value. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Further details are defined by the subject matter of the dependent claims. Hereinafter, embodiments of the present invention will be discussed in detail with reference to the drawings.

[0050] Figure 1 Schematically shown is a method sequence for processing an audio signal to improve the reproduction quality of a target part (such as the speech part of the audio signal) according to a basic embodiment;

[0051] Figure 2 Shown is a schematic flowchart illustrating an enhanced embodiment; and

[0052] Figure 3 Shown is a schematic block diagram of a decoder for processing an audio signal according to an embodiment. DETAILED DESCRIPTION

[0053] Hereinafter, embodiments of the present invention will be subsequently discussed with reference to the drawings, wherein the same reference numerals are provided to objects having the same or similar functions.

[0054] Figure 1 Shown is a schematic flowchart illustrating a method 100 including three steps / step groups 110, 120, and 130. The purpose of method 100 is to be able to process an initial audio signal AS and may have a result of outputting a modified audio signal MOD AS. The subjunctive mood is used because the possible result of outputting the audio signal MOD AS may be that there is no need to process the audio signal AS. Then, the audio signal and the modified audio signal are the same.

[0055] The three basic steps 110 and 120 are explained as step groups because here the sonar steps 110a, 110b, etc. and 120a are performed in parallel or sequentially with each other.

[0056] Within step group 110, the audio signal AS is processed individually by using different modifiers / processing methods. Here, two exemplary steps labeled 110a and 110b by reference numerals are shown, which apply a first modifier and a second modifier. These two steps can be executed in parallel or sequentially with respect to each other, and perform the processing of the audio signal AS. The audio signal can be, for example, an audio signal including one audio track, where the audio track includes two signal parts. For example, the audio track can include a speech signal part (target part) and an ambient noise signal part (side part). These two parts are labeled by reference numerals AS_TP and AS_SP. In this embodiment, it is assumed that AS_TP should be extracted from or identified within the audio signal AS in order to amplify the signal part AS_TP and thereby increase speech clarity. This process can be performed for an audio signal having only one audio track including two parts AS_SP and AS_TP, without separating the audio AS including multiple audio tracks (e.g., one audio track for AS_SP and one audio track for AS_TP).

[0057] As described above, there are multiple possible modifications of the audio signal AS, which can improve speech clarity, for example, by amplifying the AS_TP part or by reducing the AS_SP part. Other examples are reducing non-speech channels, dynamic range control, dynamic equalization, spectral sharpening, frequency shift, voice extraction, noise reduction, or other speech enhancement actions discussed in the context of the prior art. The efficiency of these modifications depends on multiple factors, for example, on the recording itself, the format of AS (e.g., a format having only one audio track or a format having multiple audio tracks), or on multiple other factors. To achieve optimal speech clarity, at least two signal modifications are applied to the signal AS. In the first step 110a, the received initial audio signal AS is modified by using a first modifier to obtain a first modified audio signal First MOD AS. Independently of step 110a, a second modification of the received initial audio signal AS is performed by using a second modifier to obtain a second modified audio signal Second MOD AS. For example, the first modifier can be based on dynamic range control, where the second modifier can be based on spectral shaping. Of course, other modifiers (e.g., based on dynamic equalization, frequency retransmission, voice extraction, noise reduction, or speech enhancement actions, or a combination of such modifiers) can also be used instead of the first modifier and / or the second modifier or as a third modifier (not shown). All methods can result in different resulting modified audio signals First MOD AS and Second MOD AS, which can be different in terms of speech clarity and similarity to the initial audio signal AS. These two parameters or at least one of these two parameters are evaluated in the next step 120.

[0058] Specifically, in step 120a, the first modified audio signal, the first MOD AS, is compared with the original audio signal, AS, in order to find the similarity. Similarly, in step 120b, the second modified audio signal, the second MOD AS, is compared with the initial audio signal, AS. For the comparison, the entity performing step 120 directly receives the audio signal AS and the first MOD AS / second MOD AS. The results of the comparison are a first perceived similarity value and a second perceived similarity value, respectively. The two values are labeled by the reference signs the first PSV and the second PSV. The two values describe the perceived similarity between the respective first modified audio signal, the first MOD AS / second modified audio signal, the second MOD AS, and the initial audio signal, AS. Assuming that the improvement in speech intelligibility is sufficient, the first modified audio signal or the second modified audio signal with the first PSV / second PSV indicating a higher similarity is selected. This is performed by the step of selection 130.

[0059] According to an embodiment, the result of the selection can be output / forwarded such that method 100 is capable of outputting the corresponding modified audio signal, the first MOD AS or the second MOD AS, having the highest similarity with the original signal. It can be seen that the modified audio signal, MOD AS, still includes two parts, AS_SP' and AS_TP'. As indicated by the (') within AS_SP' and AS_TP', both or at least one of the two parts, AS_SP' and AS_TP', are modified. For example, the magnification of AS_TP' can be increased.

[0060] According to another embodiment, an enhanced evaluation can be performed within step 120. Here, it is then further verified whether the modification performed by the first modifier or the second modifier (see steps 110a and 110b) is sufficient and improves speech intelligibility. For example, it can be analyzed where the ratio of AS_TP' to AS_SP' is greater than the ratio of AS_TP to AS_SP.

[0061] The above embodiments start from the assumption that the aim of method 100 is a MOD AS with improved speech intelligibility. According to additional embodiments, the aim of the modification can be different. For example, part of AS_TP can be another part, typically the target part that should be emphasized within the entire modified signal, MOD AS. This can be done by emphasizing / enlarging AS_TP' and / or by modifying AS_SP'.

[0062] Furthermore, the above embodiments have been discussed in the context of perceived similarity Figure 1 It should be noted that the method can be more generally used for other evaluation criteria. Figure 1Start with the assumption that the evaluation criterion is perceptual similarity. However, according to another embodiment, it is also possible to additionally use another evaluation criterion instead. For example, speech clarity can be used as the evaluation criterion. In this case, instead of step 120a, the evaluation of the first modified audio signal, the first MOD AS, is performed, and the evaluation of the second modified audio signal, the second MOD AS, is performed in step 120b. The results of these two steps of evaluations 120a and 120b are the corresponding first evaluation value and second evaluation value. After that, step 130 is performed based on the corresponding evaluation values.

[0063] Another evaluation criterion can be loudness or auditory spatial sense, etc.

[0064] Reference Figure 2 , additional embodiments with enhanced features will be discussed below.

[0065] Figure 2 A schematic flowchart showing the ability to process an audio signal AS including two parts, AS_TP (speech S) and AS_SP (ambient noise N), is shown. Here, the signal modifier 11 is used to process the signal AS such that the selection entity 13 can output a modified signal pattern AS. In this embodiment, the modifier performs different modifications 1, 2,..., M. These modifications are based on a plurality of different models, thereby generating three modified signals, the first MOD AS, the second MOD AS, and the M MOD AS. For each of the signals, the first MOD AS, the second MOD AS, and the M MOD AS, two parts, S1', N1', S2', N2', and SN', NNM' are shown. The output signals of the first MOD AS, the second MOD AS, and the M MOD AS are evaluated by the evaluator 12 for their perceptual similarity to the initial signal AS. Thus, one or more evaluator stages 12 receive the signal AS and the corresponding modified signals, the first MOD AS, the second MOD AS, and the M MOD AS. The output of this evaluation 12 is the corresponding modified signals, the first MOD AS, the second MOD AS, and the M MOD AS, and the corresponding similarity information. Based on this similarity information, the localization stage 13 determines the modulation signal MOD AS to be output.

[0066] According to an embodiment, the signal AS can be analyzed by the analyzer 21 to determine whether speech is present. In the case where there is no speech or signal to be modified within the initial audio signal AS, this decision step is marked by 21s. The initial / original audio signal AS is used as the signal, that is, unmodified (see N-MOD AS).

[0067] In the presence of speech, the second analyzer 22 analyzes whether speech clarity needs to be improved. This decision point is marked by reference numeral 22s. In the case where no modification is required, the original signal AS is used as the signal to be output (see N-MODAS). In the case where modification is suggested, the signal modifier 11 is enabled.

[0068] Based on this structure, speech clarity in audio and audiovisual media can be improved. Here, the mixed sound to be processed can be a completed mixed sound, or can include individual audio tracks or sound objects (e.g., dialogue, music, reverb, effects). In a first step, the signal is analyzed for the presence of speech (see reference numerals 21, 21s). For example, based on the mixed signal method presented in [7], the speech activity channel will be further analyzed according to physical or psychoacoustic parameters in the form of, for example, a calculated value of speech clarity (e.g., SII) or listening effort (see reference numerals 22, 22s). Based on this evaluation, by comparing the parameters with a target or threshold, it is decided whether the speech clarity is sufficient or whether sound adaptation is required. If no adaptation is required, the mixed sound proceeds as normal or remains the original mixed sound AS. If adaptation is required, an algorithm for applying a modified track or a different track to obtain the desired clarity will be applied. So far, this method is similar to the methods disclosed in US 8,195,454 B2 and US 8,271,276 B1, but is not limited to the details described in corresponding claim 1.

[0069] According to an embodiment, this means that: a model-based selection 13 of a sound reduction method that maximizes the loudness over non-speech channels (e.g., as described in US 8,577,676 B2 and US 2016 / 0071527 A1) is carried out with this concept. For the selection, another model stage 12 is applied, which simulates the perceptual similarity between the original mixed sound AS and the mixed sounds modified in different ways (first MOD AS, second MOD AS, M MOD AS) based on physical and / or psychoacoustic parameters. Here, the original mixed sound AS and different types of modified mixed sounds first MOD AS, second MOD AS, M MOD AS are used as inputs to another model stage 12.

[0070] In order to achieve the goal of best preserving the sound scene as much as possible, a method for sound adaptation can be selected (see reference numeral 13), which obtains the desired clarity through the least perceptually obvious signal modification.

[0071] According to an embodiment, possible models that can measure perceptual similarity in an instrumental manner and can be used herein are, for example, PEAQ [8], POLQA [9] or Pemo-Q

[10] . Additionally or alternatively, other physical (e.g., level) or psychoacoustic metrics (e.g., loudness, pitch, perceived source width) can be used to evaluate perceptual similarity.

[0072] An audio stream typically includes different scenarios arranged along the time domain. Thus, according to an embodiment, different sound adaptations can occur at different times in the audio track AS to have a minimally invasive perceptual effect. If, for example, the speech AS_TP and the background noise AS_TP already have significantly different spectra, a simple SNR adaptation can be the best solution because a simple SNR adaptation can best preserve the authenticity of the background noise as much as possible. If another speaker overlays the target speech, other methods (e.g., dynamic compression) may be more suitable to achieve the optimization goal.

[0073] According to a further embodiment, this model-based selection can take into account possible hearing impairments of future listeners of the audio material in the calculation, for example, in the form of an audiogram, an individual loudness function, or in the form of input personal sound preferences. Thus, speech clarity is ensured not only for people with normal hearing ability but also for people with a specific form of hearing impairment (e.g., age-related hearing loss), and it is also considered that the perceptual similarity between the original version and the processed version can vary.

[0074] Note that the analysis of speech clarity and perceptual similarity by the model and the corresponding signal processing can be performed for the entire mix or only for parts of the mix (individual scenarios, individual conversations), or can be performed in short time windows along the entire mix, such that a decision can be made for each window as to whether a sound adaptation must be performed.

[0075] Below, an example of such a process will be discussed exemplarily:

[0076] i. No sound adaptation: If the analysis of the listening model indicates that sufficient high speech clarity is ensured, no further sound adaptation will be performed. Alternatively, the following adaptation is performed to avoid perceptual differences between different scenarios. An "interpolation" between no processing and the processing selected below can also be performed. The two modes can achieve perceptual continuity on different time frames / scenarios.

[0077] For the separated tracks of dialogue and background noise, the following steps can be performed:

[0078] ii. Adapt the sound signal: For example, by increasing the level, by frequency weighting and / or single-channel or multi-channel dynamic compression, only the track of the speech signal is processed to improve speech clarity.

[0079] iii. Adapt interference noise: Process one or several audio tracks that do not include speech, for example, by reducing the level, by frequency weighting, and / or by single-channel or multi-channel dynamic compression, to improve speech intelligibility. However, for reasons of sound aesthetics, the simple case of completely eliminating background noise resulting in improved speech intelligibility is not practical because the design of music, effects, etc. is also an important part of creative sound design.

[0080] iv. Adapt all tracks: One or several of the tracks of the speech signal and other tracks are processed by the above methods to improve speech intelligibility.

[0081] Note that for adaptation, artificial intelligence such as using neural networks can be used. In the already mixed audio signal (i.e., the non-separated tracks of dialogue and background noise), for example, when a source separation method is pre-used, steps ii to iv can also be performed. This source separation method separates the mixed sound into speech and one or several background noises. Then, improving speech intelligibility can include, for example, remixing the separated signals with an improved SNR, or modifying the speech signal and / or background noise or part of the background noise by frequency weighting or single-channel or multi-channel dynamic compression. Here, the sound adaptation that both improves speech intelligibility as desired and at the same time best preserves the original sound as possible will be selected again. The method for source separation can be applied without any explicit stage for detecting speech activity.

[0082] Note that according to an embodiment, the selection of the corresponding processing can be performed by using artificial intelligence / neural networks. For example, if there are more than one factor for selection (such as a perceived value and a loudness value or a value describing a match with personal listening preferences), then the artificial intelligence / neural network can be used.

[0083] The adaptation of the scenario (even if this is not necessary) to maintain perceptual continuity over different time frames / scenarios has been discussed above. According to another variant, the adaptation of multiple or all scenarios can be selected. In addition, it should be noted that between different scenarios, a transition between different adaptation scenarios or between an adaptation scenario and a non-adaptation scenario can be integrated to maintain perceptual continuity.

[0084] According to an embodiment, the assessment and optimization based on perceptual similarity (see reference numeral 12) may relate to the target language, background noise, or a mixture of speech and background noise. There may be different thresholds, for example, for processing the perceptual similarity of a speech signal, processing background noise, or processing a mixed sound with the corresponding original signal, such that a specific degree of signal modification of the corresponding signal may not be exceeded. Another boundary condition may be that the background noise (e.g., music) may not change perceptually too much relative to a previous or subsequent time point, because otherwise, when, for example, speech is present, the perceptual continuity will be disturbed, the music will be reduced too much or changed in its frequency content, or the speech of an actor may not change too much during the course of a movie. Such boundary conditions may also be checked based on the above model.

[0085] This may have the following effect: Without overly disturbing the perceptual similarity of the speech and / or background noise, it may not be possible to obtain the desired improvement in clarity. Here, a (possibly configurable) decision-making phase may decide which goal is to be achieved, or whether and how to find a compromise.

[0086] Here, the processing may be carried out iteratively, i.e., the listening model may be checked again after sound adaptation to verify that the desired speech clarity and perceptual similarity to the original speech have been obtained.

[0087] The processing may be carried out over the entire duration of the audio material or only over a part of the duration of the audio material (e.g., a scene, a dialogue), depending on the calculation of the listening model.

[0088] The embodiment may be used for all audio and audiovisual media (movies, radio, podcasts, general audio rendering). Possible commercial applications are, for example:

[0089] i. Internet-based services, where a customer loads his audio material, activates automatic speech clarity improvement, and downloads the processed signal. Internet-based services may be extended by customer-specific selection of the sound adaptation method and the degree of sound adaptation. Such services already exist, but do not use a listening model for sound adaptation regarding speech clarity (see above 2. (V.)).

[0090] ii. Software solutions for sound production tools, for example, integrated in a digital audio workstation (DAW) to enable correction of archived or current mixes.

[0091] iii. Test algorithms that identify channels in the audio material that do not correspond to the desired speech clarity and may provide the user with suggested sound adaptation modifications for selection.

[0092] iv. Software and / or hardware, integrated in a terminal device at the listener side of a broadcast chain, such as a soundbar, headphones, a television device, or a device receiving streaming audio content.

[0093] In Figure 1 the context of the method discussed in Figure 2 or the concept discussed in Figure 3 can be implemented by using a processor. The processor is shown by

[0094] Figure 3 The processor 10 in the two-stage signal modifier 11 and the evaluator / selector 12 and 13 is shown. The modifier receives an audio signal from an interface and performs modifications based on different models in order to obtain a modified audio signal MOD AS. The evaluator / selector 12 receives an audio signal from the interface and performs modifications based on different models in order to obtain a modified audio signal MOD AS. The evaluator / selector 12, 13 evaluates the similarity and selects, based on this information, the signal with the highest similarity or with high similarity and improved speech intelligibility (which is sufficient) to output MOD AS.

[0095] Of course, the two stages 11, 12 and 13 can be implemented by one processor.

[0096] Although some aspects have been described in the context of a device, it will be clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of the features of the corresponding block or item or the corresponding device. Some or all of the method steps can be performed by (or using) a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0097] The novel encoded audio signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (e.g., the Internet).

[0098] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or in software. Implementations can be performed using a digital storage medium storing electronically readable control signals (e.g., a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory) which cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Thus, the digital storage medium can be computer-readable.

[0099] Some embodiments according to the invention include a data carrier having an electronically readable control signal, which is capable of cooperating with a programmable computer system in order to execute one of the methods described herein.

[0100] Generally, embodiments of the invention may be implemented as a computer program product having program code operable to execute one of the methods when the computer program product is run on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0101] Other embodiments include a computer program stored on a machine-readable carrier for executing one of the methods described herein.

[0102] In other words, embodiments of the method according to the invention are thus computer programs having program code for executing one of the methods described herein when the computer program is run on a computer.

[0103] Thus, another embodiment of the method according to the invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded for executing one of the methods described herein. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory.

[0104] Thus, another embodiment of the method according to the invention is a data stream or signal sequence representing a computer program for executing one of the methods described herein. The data stream or signal sequence may be configured, for example, to be transmitted via a data communication connection (e.g., via the Internet).

[0105] Another embodiment includes a processing device, e.g., a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.

[0106] Another embodiment includes a computer on which a computer program is installed for executing one of the methods described herein.

[0107] Another embodiment according to the invention includes a device or system configured to transmit (e.g., electronically or optically) a computer program to a receiver, the computer program for executing one of the methods described herein. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The device or system may include, for example, a file server for transmitting the computer program to the receiver.

[0108] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to execute some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to execute one of the methods described herein. Generally, the method is preferably executed by any hardware device.

[0109] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to other technicians in the art. Therefore, it is intended to be limited only by the scope of the appended patent claims rather than by the specific details given by the description and interpretation of the embodiments herein.

[0110] References

[0111] [1] Simon, C. and Fassio, G. (2012). Optimierung audiovisuellerMedien für Hörgeschädigte. In: Fortschritte der Akustik – DAGA 2012,Darmstadt, March 2012.

[0112] [2] Ephraim, Y. und Malah, D. (1984). Speech enhancement using aminimum-mean square error short-time spectral amplitude estimator. IEEETransactions on Acoustics Speech and Signal Processing, 32(6):1109–1121.

[0113] [3] Kolbæk, M., Yu, D., Tan, Z-H., & Jensen, J. (2017). MultitalkerSpeech Separation With Utterance-Level Permutation Invariant Training of DeepRecurrent Neural Networks. IEEE Transactions on Audio, Speech and LanguageProcessing, 25(10), 1901-1913. https: / / doi.org / 10.1109 / TASLP.2017.2726762

[0114] [4] Jouni, P., Torcoli, M., Uhle, C., Herre, J., Disch, S., Fuchs, H.(2019). Source Separation for Enabling Dialogue Enhancement in Object-based Broadcast with MPEG-H. JAES 67, 510-521. https: / / doi.org / 10.17743 / jaes.2019.0032

[0115] [5] Sauert, B. and Vary, P. (2012). Near end listening enhancement in the presence of bandpass noises. In: Proc. der ITG-Fachtagung Sprachkommunikation, Braunschweig, September 2012.

[0116] [6] ANSI S3.5 (1997). Methods for calculation of speech intelligibility index.

[0117] [7] Huber, R., Pusch, A., Moritz, N., Rennies, J., Schepker, H., Meyer, B.T. (2018). Objective Assessment of a Speech Enhancement Scheme with an Automatic Speech Recognition-Based System. ITG-Fachbericht 282: Speech Communication, 10. – 12. October 2018 in Oldenburg, 86-90.

[0118] [8] ITU-R Recommendation BS.1387: Method for objective measurements of perceived audio quality (PEAQ)

[0119] [9] ITU-T Recommendation P.863: Perceptual objective listening quality assessment

[0120]

[10] Huber, R. and Kollmeier, B. (2006). PEMO-Q—A New Method for Objective Audio Quality Assessment Using a Model of Auditory Perception. IEEE Transactions on Audio, Speech, and Language Processing 14(6), 1902-1911

[0121]

[11] NetMix player of Fraunhofer IIS, http: / / www.iis.fraunhofer.de / de / bf / amm / forschundentw / forschaudiomulti / dialogenhanc.html

[0122]

[12] https: / / auphonic.com / 。

Claims

1. A method (100) for processing an initial audio signal comprising a target portion and a side portion, comprising the steps of: a. Receiving the initial audio signal; b. Modifying (110, 110a) the received initial audio signal by using a first signal modifier to obtain a first modified audio signal; Modifying (110, 110b) the received initial audio signal by using a second signal modifier to obtain a second modified audio signal; c. Evaluating (120, 120a) the first modified audio signal against an evaluation criterion to obtain a first evaluation value describing the degree of satisfaction of the evaluation criterion; Evaluating (120, 120a) the second modified audio signal against the evaluation criterion to obtain a second evaluation value describing the degree of satisfaction of the evaluation criterion; And d. Selecting (130) the first modified audio signal or the second modified audio signal depending on the respective first evaluation value or second evaluation value; wherein the selection step is performed based on a plurality of independent first evaluation values and independent second evaluation values, or based on at least two independent evaluation criteria, wherein the evaluation criterion is perceptual similarity, and wherein step c comprises the following sub-steps: Comparing (120, 120a) the received initial audio signal with the first modified audio signal to obtain a first perceptual similarity value as the first evaluation value, the first perceptual similarity value describing the perceptual similarity between the initial audio signal and the first modified audio signal; and Comparing (120, 120b) the received initial audio signal with the second modified audio signal to obtain a second perceptual similarity value as the second evaluation value, the second perceptual similarity value describing the perceptual similarity between the initial audio signal and the second modified audio signal.

2. The method (100) according to claim 1, wherein, The evaluation criterion is from the group comprising: - Perceptual similarity, described by a first perceptual similarity value and a second perceptual similarity value, the first perceptual similarity value and the second perceptual similarity value describing the perceptual similarity between the respective first modified audio signal and second modified audio signal and the initial audio signal; - Speech intelligibility, compared with a target or threshold in the form of a calculated value of speech intelligibility; - Loudness, described by a loudness value; - Sound pattern; - Spatial sense.

3. The method (100) according to claim 1, wherein, The at least two independent evaluation criteria are evaluated separately such that respective first evaluation values describing the degree of satisfaction of the at least two independent evaluation criteria for the first modified audio signal and respective second evaluation values describing the degree of satisfaction of the at least two independent evaluation criteria for the second modified audio signal are determined, wherein the selection is then performed based on the weighted first evaluation values and weighted second evaluation values.

4. The method (100) according to claim 1, wherein, Selecting the first modified audio signal, wherein the first perceptual similarity value is higher than the second perceptual similarity value to indicate a higher perceptual similarity of the first modified audio signal; and Wherein, when the second perceptual similarity value is higher than the first perceptual similarity value to indicate a higher perceptual similarity of the second modified audio signal, the second modified audio signal is selected.

5. The method (100) according to claim 1 further comprises the following steps: Output the first modified audio signal or the second modified audio signal depending on the selection in step d.

6. The method (100) according to claim 3, wherein, When the corresponding first perceptual similarity value or second perceptual similarity value is lower than a threshold, perform the step of outputting the original audio signal instead of outputting the first modified audio signal or the second modified audio signal, wherein below the threshold, the corresponding first modified audio signal or second modified audio signal is indicated as not being similar enough to the original audio signal.

7. The method (100) according to claim 1, wherein, The target part is the speech part of the original audio signal, and the side part is the ambient noise part of the original audio signal.

8. The method (100) according to claim 1, wherein, The first modified audio signal and / or the second modified audio signal includes: the target part moved to the foreground and the side part moved to the background, and / or the speech part moved to the foreground as the target part and the ambient noise part moved to the background as the side part.

9. The method (100) according to claim 1, wherein The step of comparison includes: extracting the first evaluation value and / or the second evaluation value by using a perceptual model, a PEAQ model, a POLQA model, and / or a PEMO-Q model.

10. The method (100) according to claim 1, wherein, The first evaluation value and / or the second evaluation value depends on the physical parameters of the first modified audio signal or the second modified audio signal, the volume level of the first modified audio signal or the second modified audio signal, the psychoacoustic parameters of the first modified audio signal or the second modified audio signal, the loudness information of the first modified audio signal or the second modified audio signal, the pitch information of the first modified audio signal or the second modified audio signal, and / or the perceptual source width information of the first modified audio signal or the second modified audio signal.

11. The method (100) according to claim 1, wherein, The first signal modifier and / or the second signal modifier is configured to perform SNR increase, dynamic compression, SNR increase of the original audio signal, and / or dynamic compression of the original audio signal; and / or wherein, if the original audio signal includes a separate target part and a separate side part, the step of modification includes: increasing the target part, increasing the frequency weighting of the target part, dynamically compressing the target part, reducing the side part, reducing the frequency weighting of the side part; and / or wherein, if the original audio signal includes a combined target part and side part, the modification includes: performing separation of the target part and the side part.

12. The method (100) according to claim 1, wherein, The step of selection (130) is performed in consideration of one or more of the following factors: - The hearing loss level of the hearing-impaired person; - Individual hearing performance; - Individual frequency-related hearing performance; - Individual preference; - Individual preference regarding the signal modification rate.

13. The method (100) according to claim 1, wherein, The step of modification (110) and / or comparison (120) is performed in consideration of one or more of the following factors: The hearing loss level of the hearing-impaired person; Individual hearing performance; Individual frequency-related hearing performance; Individual preferences; Individual preferences regarding the rate of signal modification.

14. The method (100) according to claim 1, wherein, The method further comprises the steps of: receiving information regarding an optimization goal that defines an individual preference; wherein the evaluation criterion depends on the optimization goal; or wherein the steps of modifying and / or evaluating and / or selecting depend on the optimization goal; or wherein the weighting of an independent first evaluation value and a second evaluation value of an independent evaluation criterion that is independent of the description of the selection step depends on the optimization goal.

15. The method (100) according to claim 1, wherein, The step of comparison (120) is performed for: the entire initial audio signal with the entire first modified audio signal and the second modified audio signal; and / or Target portions of individual audio signals with corresponding target portions of the first modified audio signal and the second modified audio signal; and / or Side portions of the initial audio signal with side portions of the first modified audio portion and the second modified audio portion.

16. The method (100) according to claim 1, wherein The initial audio signal comprises a plurality of time frames, and wherein steps a to d are repeated for each time frame; and / or wherein steps a to d are repeated for a temporal portion or a time frame of a scene of the initial audio signal.

17. The method (100) according to claim 1, wherein, The adaptation of the initial audio signal comprising a plurality of time frames is performed for the time frames that require the adaptation and other time frames in order to maintain perceptual continuity, or wherein the adaptation of the initial audio signal comprising a plurality of time frames is performed for the time frames that require the adaptation and in an interpolated manner for other time frames in order to maintain perceptual continuity; and / or wherein the adaptation of a first subsequent time frame and a second subsequent time frame is performed such that a transition between the first subsequent time frame and the second subsequent time frame is formed in order to maintain perceptual continuity.

18. The method (100) according to claim 1, wherein, The method (100) further comprises the following initial steps: Analyzing (21) an initial audio portion to determine a speech portion; Comparing the speech portion with an ambient noise portion in order to evaluate the speech intelligibility of the initial audio signal; And Activating the first signal modifier and / or the second signal modifier for modification if a value indicative of the speech intelligibility is below a threshold.

19. A computer-readable storage medium having stored thereon a computer program with program code for performing the method according to claim 1 when run on a computer.

20. An apparatus for processing an initial audio signal comprising a target portion and a side portion, the apparatus comprising: An interface for receiving the initial audio signal; A first signal modifier (11) and a second signal modifier (11), the first signal modifier (11) for modifying (110) the received initial audio signal to obtain a first modified audio signal, the second signal modifier (11) for modifying the received initial audio signal to obtain a second modified audio signal; An evaluator for evaluating (120, 120a) the first modified audio signal against an evaluation criterion to obtain a first evaluation value describing the degree of satisfaction of the evaluation criterion, and for evaluating (120, 120a) the second modified audio signal against the evaluation criterion to obtain a second evaluation value describing the degree of satisfaction of the evaluation criterion; and A selector (13) for selecting (130) the first modified audio signal or the second modified audio signal depending on the respective first evaluation value or second evaluation value; wherein the selection step is performed based on a plurality of independent first evaluation values and independent second evaluation values, or based on at least two independent evaluation criteria; wherein the evaluation criterion is perceptual similarity, and wherein the evaluator is configured, when evaluating: to compare (120, 120a) the received initial audio signal with the first modified audio signal to obtain a first perceptual similarity value as the first evaluation value, the first perceptual similarity value describing the perceptual similarity between the initial audio signal and the first modified audio signal; and to compare (120, 120b) the received initial audio signal with the second modified audio signal to obtain a second perceptual similarity value as the second evaluation value, the second perceptual similarity value describing the perceptual similarity between the initial audio signal and the second modified audio signal.

Citation Information

Patent Citations

  • Method and System for Scaling Ducking of Speech-Relevant Channels in Multi-Channel Audio

    US20160071527A1

  • Speech enhancement in entertainment audio

    US8195454B2

  • Enhancement of multichannel audio

    US8271276B1

  • Method and apparatus for maintaining speech audibility in multi-channel audio with minimal impact on surround experience

    US8577676B2

  • A speech intelligibility predictor and applications thereof

    CN102194460A