Apparatus and method for audio signal processing to advantageously modify coherent portions of audio signal

By separating and phase-aligning the coherent signal portions in a compact audio device, the problem of signal cancellation in multi-channel signal reproduction is solved, achieving effective reproduction of all content, especially the preservation of inverted signals.

CN121844580APending Publication Date: 2026-04-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

When compact audio devices process multi-channel signals, there is a problem where some signals are canceled out, making them unreproducible, especially the inverted signal portion, which cannot be effectively reproduced on a single speaker.

Method used

The audio input signal is separated into coherent and incoherent parts by a signal splitter, the coherent part is phase-aligned by a signal processor, and the audio output signal is generated by a combiner to avoid signal cancellation.

Benefits of technology

Ensuring that all audio content can be reproduced, especially that the inverted signal portion is not eliminated, achieves more accurate spatial audio reproduction on devices with limited loudspeakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844580A_ABST
    Figure CN121844580A_ABST
Patent Text Reader

Abstract

An apparatus for audio signal processing according to an embodiment is provided. The apparatus comprises a signal separator (110) for separating each of the at least two audio input signals into a first signal portion and a second signal portion. Furthermore, the apparatus comprises a signal processor (120) for obtaining, from a first signal portion of each of the at least two audio input signals, a phase-aligned signal portion of each of the at least two audio input signals by modifying the first signal portion of at least one of the at least two audio input signals; wherein the signal processor (120) is configured to modify a first signal portion of the at least one audio input signal by phase aligning the first signal portion of the at least one audio input signal with a first signal portion of at least a further one of the at least two audio input signals. Furthermore, the apparatus comprises a combiner (130) for combining the phase-aligned signal portion and the second signal portion of each of the at least two audio input signals to obtain at least two audio output signals.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to audio processing, to an apparatus and method for audio signal processing to advantageously modify the coherent part of an audio signal, and more specifically to a (pre-) processing to advantageously modify the coherent part of an audio signal. BACKGROUND

[0002] In recent years, compact audio devices like soundbars and smart speakers have become increasingly popular. In contrast to traditional speaker setups, where dedicated speakers are used to reproduce the content of a single input channel, these compact reproduction devices typically only have a limited number of speakers ("limited number of speakers" can for example mean "single device with a limited number of drivers"). The simplest smart speaker consists of only a single full-range driver for audio playback.

[0003] In order to be able to reproduce or at least simulate the reproduction of the spatial impression intended by the original content signal, smart speakers or soundbars with more than one speaker driver typically include spatial audio processing, which creates a spatial impression using acoustical or psycho-acoustical means.

[0004] The most common type of input signal in today's consumer environment is still two-channel stereo content, while the number of surround content (e.g. 5.1 or 7.1) and immersive content with height channels (e.g. 5.1+4 or 7.1+4), different orders of surround signals and object-based audio content is continuously increasing.

[0005] In order to be able to reproduce such multi-channel signals on the above-mentioned consumer playback devices, the signals of the different channels need to be combined at some point in the processing, before being reproduced by the limited number of speakers.

[0006] During content creation, specific perceptual effects are evoked by audio recording, mixing and rendering using gain differences, delay differences and phase differences between signal components on different channels or objects.

[0007] If such content is reproduced using the compact consumer devices instead of the intended playback setup, the combination of these signals for playback on the compact device can lead to a deviation from the original signal and to a deviation from the evoked perception.

[0008] One of the most critical cases that can occur (and which is prevented by the method of the invention) leads to the complete cancellation of signal content (which can be the complete signal or only parts or components of the signal, depending on the specific case of the content), which means that they will be completely inaudible, which would be a drastic change of the content.

[0009] In the following example illustration (use case description) we exemplarily use the simplest reproduction device as an example, which consists of a single single-channel smart loudspeaker with only a single loudspeaker driver, which is fed by a two-channel input signal.

[0010] Figure 2 A device-specific processing is shown. In this scenario, two input signals are combined for playback through a single driver. In this case, signal cancellation occurs when two input signals carry anti-phase signals or signals with anti-phase portions, which will be cancelled when combined for playback through a single driver. In this way, the signal content carried by the anti-phase signal portions will be lost in the reproduction.

[0011] This situation is not ideal, as anti-phase signal portions are often included in productions for specific reasons. One of these is to create a specific perceptual effect when two anti-phase signals are played back through two separate loudspeakers.

[0012] While this effect cannot be achieved by playing back the signals on only a single loudspeaker, it is still desirable to preserve the content of these signals in the reproduced sound. SUMMARY

[0013] The method of the invention described below avoids the loss of this signal portion, so that all content is audible.

[0014] Figure 3 A second scenario is shown, in which a device with multi-channel input, two loudspeakers and spatial processing is considered, which employs dipole processing (also known as gradient processing). The purpose of gradient processing is to invert the phase of the signals when applied to multiple loudspeakers in order to generate a specific directivity pattern of the playback device.

[0015] In Figure 3 , the inputs In_1 to In_5 can correspond to the left channel, center channel, right channel, left surround channel and right surround channel of a 5.0 surround sound signal, respectively.

[0016] The left signal (In_1) is reproduced by the left driver of the device.

[0017] The right channel (In_3) is reproduced by the right driver of the device.

[0018] The center channel (In_2) is split and reproduced by the two loudspeakers of the device.

[0019] The surround channels (In_4 and In_5) are fed to the two loudspeakers of the device in a dipole fashion. This is indicated by applying a phase inversion (multiplication by -1) to the split signal, which is fed to one of the drivers. Note that for these two signals, the phase inversion is applied to different loudspeakers.

[0020] The processing of the input channels in such a device usually comprises more steps and is more complex. For example, the center signal will be attenuated to avoid being more prominent than the left and right signals when played back through both loudspeakers.

[0021] Furthermore, additional processing can be applied for the surround channels and the dipole processing can have further parameters like gain and delay applied to both input signals to control the achieved directional effect.

[0022] For the purpose of example, Figure 3 The core is highlighted which is the phase inversion (multiplication by -1) applied to the signals. Similar processing is also applied in devices with more than two loudspeakers to achieve directional reproduction or specific directivity patterns.

[0023] Different methods and implementations are known in the literature.

[0024] If the input signals of such a differential processing carry positively correlated (see below) signal components, these components will be cancelled out when played back. In the given example, one example where this happens is when a certain signal is positioned between two surround channels.

[0025] It is an object of the present invention to provide an improved audio signal processing concept. The object of the present invention is achieved by the apparatus according to claim 1, the method according to claim 65 and the computer program according to claim 66.

[0026] According to an embodiment, an apparatus for audio signal processing is provided. The apparatus comprises a signal separator for separating each of at least two audio input signals into a first signal part and a second signal part. Furthermore, the apparatus comprises a signal processor for deriving a phase-aligned signal part of each of the at least two audio input signals from the first signal part of each of the at least two audio input signals by modifying the first signal part of at least one of the at least two audio input signals; wherein the signal processor is configured to modify the first signal part of the at least one of the at least two audio input signals by phase-aligning the first signal part of the at least one of the at least two audio input signals with the first signal part of at least one further of the at least two audio input signals. Furthermore, the apparatus comprises a combiner for combining the phase-aligned signal part with the second signal part of each of the at least two audio input signals to derive at least two audio output signals.

[0027] Furthermore, according to an embodiment, a method for audio signal processing is provided. The method comprises:

[0028] - separating each of the at least two audio input signals into a first signal part and a second signal part;

[0029] - obtaining a phase-aligned signal part of each of the at least two audio input signals from the first signal part of each of the at least two audio input signals by modifying the first signal part of at least one of the at least two audio input signals; wherein modifying the first signal part of said at least one audio input signal is achieved by phase-aligning the first signal part of said at least one audio input signal with the first signal part of at least one other of the at least two audio input signals; and

[0030] - combining the phase-aligned signal part and the second signal part of each of the at least two audio input signals to obtain at least two audio output signals.

[0031] Further, a computer program is provided according to an embodiment, which, when executed on a computer or signal processor, is configured to implement the above described method.

[0032] Some embodiments relate to a processor that processes an audio input signal in such a way that it is adjusted in a certain way to avoid adverse effects that can occur in subsequent processing.

[0033] Preferred embodiments relate to the field of audio reproduction. Although the following is described as an example application in the context of reproduction, the processing can also be applied in other contexts, such as content production, audio encoding, audio signal transmission, etc.

[0034] The embodiments avoid the loss of positively correlated signal parts that would be cancelled out by the differential processing in prior art systems' reproduction. BRIEF DESCRIPTION OF DRAWINGS

[0035] In the following, embodiments of the present application will be described in more detail with reference to the accompanying drawings, in which:

[0036] Figure 1 An apparatus for audio signal processing according to an embodiment is shown.

[0037] Figure 2 A device-specific processing is shown, in which two input signals are combined for playback on a single driver.

[0038] Figure 3 A second scenario is shown, in which a device with multi-channel input, two loudspeakers, and spatial processing is considered, which employs dipole processing.

[0039] Figure 4A scenario is shown in which a processor receives two audio signals at its input, processes the two audio signals, and outputs two audio signals.

[0040] Figure 5 More details of audio signal processing according to embodiments are shown.

[0041] Figure 6 A diagram is shown according to embodiments in which the weighting factor decreases with increasing frequency.

[0042] Figure 7 Separate functions for different threshold values are shown according to embodiments.

[0043] Figure 8 A plot is shown depicting an example mapping between a correlation indicator and a correlation adaptation time constant according to embodiments.

[0044] Figure 9 Examples of smoothing attack time and release time according to embodiments are shown.

[0045] Figure 10 An example application with smart speakers according to embodiments is shown.

[0046] Figure 11 A scenario is shown according to embodiments in which audio signals are transmitted from a source device to multiple playback devices.

[0047] Figure 12 A device is shown according to embodiments with two loudspeaker drivers in a single enclosure.

[0048] Figure 13 An example application of processing in a soundbar device according to embodiments is shown.

[0049] Figure 14 An embodiment of multiple input channels applying processing according to embodiments multiple times in a parallel fashion is shown.

[0050] Figure 15 An embodiment of multiple input channels applying processing according to embodiments multiple times in a serial / sequential fashion is shown.

[0051] Figure 16 An embodiment of multiple input channels applying processing according to embodiments multiple times by combining Figure 14 a parallel fashion and Figure 15 a serial / sequential fashion is shown.

[0052] Figure 17Embodiments of a multi-input channel by extending the processor to support multiple inputs and adding means to select the two input channels that should be processed based on a selection parameter or a control parameter are schematically shown.

[0053] Figure 18 Embodiments of a multi-input channel by calculating the coherence and correlation between multiple channels and different combinations thereof and modifying the processor to have the phase alignment occur from one channel to multiple other channels are schematically shown.

[0054] Figure 19 Embodiments without power compensation are shown.

[0055] Figure 20 A transfer function according to an embodiment is shown. DETAILED DESCRIPTION

[0056] Figure 1 An apparatus for audio signal processing according to an embodiment is shown.

[0057] The apparatus comprises a signal separator 110 for separating each of the at least two audio input signals into a first signal part and a second signal part.

[0058] Further, the apparatus comprises a signal processor 120 for obtaining a phase-aligned signal part of each of the at least two audio input signals from the first signal part of each of the at least two audio input signals by modifying the first signal part of at least one of the at least two audio input signals; wherein the signal processor 120 is configured to modify the first signal part of the at least one audio input signal by phase-aligning the first signal part of the at least one audio input signal with the first signal part of at least one further audio input signal of the at least two audio input signals.

[0059] Further, the apparatus comprises a combiner 130 for combining the phase-aligned signal part with the second signal part of each of the at least two audio input signals to obtain the at least two audio output signals.

[0060] According to an embodiment, the signal separator 110 can for example be configured to separate each of the at least two audio input signals into the first signal part and the second signal part depending on the coherence and / or the correlation.

[0061] In an embodiment, the signal separator 110 can for example be configured to separate each of the at least two audio input signals into the first signal part and the second signal part depending on the coherence and / or the correlation of the first signal part with a signal part of one or more other audio input signals of the at least two audio input signals.

[0062] According to an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal portion (e.g. a coherent signal portion) and a second signal portion (e.g. a non-coherent signal portion), such that the first signal portion can be coherent with signal portions of one or more other audio input signals of the at least two audio input signals, for example.

[0063] In an embodiment, in order to obtain the phase-aligned signal portion of each of the at least two audio input signals, the signal processor 120 can be configured to modify the first signal portion of the at least one audio input signal and configured not to modify the first signal portion of the at least one other audio input signal, for example.

[0064] According to an embodiment, in order to obtain the phase-aligned signal portion of each of the at least two audio input signals, the signal processor 120 can be configured to modify the first signal portion of each of the at least two audio input signals, for example.

[0065] In an embodiment, the at least two audio input signals can be exactly two audio input signals, for example, the at least one audio input signal can be exactly one audio input signal, for example, the at least one other audio input signal can be exactly one other audio input signal, for example, and the at least two audio output signals can be exactly two audio output signals, for example.

[0066] According to an embodiment, the signal processor 120 can be configured to phase-align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal in the frequency domain, for example.

[0067] In an embodiment, the signal processor 120 can be configured to align the phase of at least one frequency band of the first signal portion of the at least one audio input signal with the phase of at least one frequency band of the first signal portion of the at least one other audio input signal in the frequency domain, for example.

[0068] According to an embodiment, the signal processor 120 can be configured to align the phase of each of two or more frequency bands of the first signal portion of the at least one audio input signal with the phase of each of the two or more other frequency bands of the first signal portion of the at least one other audio input signal in the frequency domain, for example.

[0069] In an embodiment, the apparatus further comprises a time-to-frequency transform unit for transforming the at least two audio input signals represented in the time domain from the time domain to the frequency domain. The apparatus further comprises a frequency-to-time transform unit for transforming the at least two audio output signals represented in the frequency domain from the frequency domain to the time domain.

[0070] According to an embodiment, the time-to-frequency transform unit can for example be configured to perform a short-time Fourier transform to transform the at least two audio input signals from the time domain to the frequency domain. The frequency-to-time transform unit can for example be configured to perform a short-time inverse Fourier transform to transform the at least two audio output signals from the frequency domain to the time domain.

[0071] In an embodiment, the signal processor 120 can for example be configured to phase align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal such that, in case the first signal portion of the at least one audio input signal is negatively correlated with the first signal portion of the at least one other audio input signal, the phase aligned signal portion of the at least one audio input signal and the phase aligned signal portion of the at least one other audio input signal have the same phase after the phase alignment.

[0072] According to an embodiment, the signal processor 120 can for example be configured to phase align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal such that, in case the first signal portion of the at least one audio input signal is positively correlated with the first signal portion of the at least one other audio input signal, the phase aligned signal portion of the at least one audio input signal and the phase aligned signal portion of the at least one other audio input signal have inverted phases after the phase alignment.

[0073] In an embodiment, the signal processor 120 can for example be configured to phase align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal by copying phase information of the first signal portion of the at least one other audio input signal to the first signal portion of the at least one audio input signal.

[0074] According to an embodiment, the signal processor 120 can for example be configured to phase align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal by copying and inverting phase information of the first signal portion of the at least one other audio input signal to the first signal portion of the at least one audio input signal.

[0075] In an embodiment, the signal processor 120 can for example be configured to phase align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal without changing the amplitude of the first signal portion of the at least one audio input signal and without changing the amplitude of the first signal portion of the at least one other audio input signal.

[0076] According to an embodiment, the second signal portion of each of the at least two audio input signals can be unmodified, e.g. when being combined by the combiner 130.

[0077] In an embodiment, the at least two audio input signals can comprise, e.g., one or more audio channel signals and / or one or more audio object signals and / or one or more ambisonic signals.

[0078] In an embodiment, the apparatus comprises a power compensator such that a total signal energy of the at least two audio output signals corresponds to a total signal energy of the at least two audio input signals, or such that a signal energy of one of the at least two audio output signals corresponds to a signal energy of one of the at least two audio input signals, or such that a signal energy of each of the at least two audio output signals corresponds to a signal energy of one of the at least two audio input signals.

[0079] According to an embodiment, the power compensator can be configured to perform power compensation per frequency bin or per frequency band, e.g.

[0080] In an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal portion and a second signal portion by applying a first mask value to a time-frequency bin of the audio input signal to obtain a time-frequency bin of the first signal portion, and by applying a second mask value to the time-frequency bin of the audio input signal which depends on the first mask value to obtain a time-frequency bin of the second signal portion, e.g.

[0081] According to an embodiment, the signal separator 110 can be configured to apply the same first mask value to two or more time-frequency bins of the same frequency band of the audio input signal to obtain two or more time-frequency bins of the first signal portion of the same frequency band, and / or the signal separator 110 can be configured to apply the same second mask value to two or more time-frequency bins of the same frequency band of the audio input signal to obtain two or more time-frequency bins of the second signal portion of the same frequency band, e.g.

[0082] In an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal portion and a second signal portion by multiplying a first mask value with the time-frequency bin of the audio input signal to obtain a time-frequency bin of the first signal portion, and by multiplying a second mask value with the time-frequency bin of the audio input signal to obtain a time-frequency bin of the second signal portion, wherein the first mask value assumes a value vi with 0 < vi < 1, and wherein the second mask value v2 = 1 - vi, e.g.

[0083] According to an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal part and a second signal part depending on a coherence of a time-frequency bin of the plurality of time-frequency bins.

[0084] In an embodiment, the signal separator 110 can be configured to update the first mask value and the second mask value for each of the plurality of time-frequency bins of the audio input signals such that the first signal part comprises only the coherent signal parts of the at least two audio input signals whose sum exhibits a potential cancellation greater than a threshold value.

[0085] According to an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal part and a second signal part depending on a coherence of each time-frequency bin of the plurality of time-frequency bins, wherein the coherence is averaged over time.

[0086] In an embodiment, the signal separator 110 can be configured to determine the coherence of each time-frequency bin of the plurality of time-frequency bins depending on an autocorrelation of the time-frequency bins averaged over time and depending on a cross-correlation of the time-frequency bins averaged over time.

[0087] According to an embodiment, the signal separator 110 can be configured to determine a frequency-dependent absolute cross-spectral phase, which is summarized to a single absolute cross-spectral phase value by a frequency-dependent mean value.

[0088] In an embodiment, a single frequency-dependent absolute cross-spectral phase value exhibiting a value of 0 indicates a positive correlation, a single frequency-dependent absolute cross-spectral phase value exhibiting a value of 0.5 indicates no correlation, and a single frequency-dependent absolute cross-spectral phase value exhibiting a value of 1 indicates a negative correlation.

[0089] According to an embodiment, the signal separator 110 can be configured to separate each of the at least two audio input signals into a first signal part and a second signal part by employing a separation function, which depends on a coherence of a time-frequency bin of the plurality of time-frequency bins.

[0090] In an embodiment, the separation function separates an amplitude of the time-frequency bin into a coherent amplitude part and an incoherent amplitude part.

[0091] According to an embodiment, the separation function can be frequency-dependent, for example.

[0092] In an embodiment, the separation function depends on a signal property of at least one of the at least two audio input signals.

[0093] According to an embodiment, the separation function depends on the threshold value.

[0094] In an embodiment, the threshold value can for example be frequency dependent, such that the signal separator 110 can for example be configured to assign a greater amplitude portion to a first signal portion that exhibits a lower frequency than to an amplitude portion that is assigned to a first signal portion that exhibits a higher frequency for the same coherence.

[0095] According to an embodiment, the apparatus comprises an interface for setting the threshold value.

[0096] In an embodiment, the interface can for example be configured to set the threshold value separately per frequency band or to set the threshold value separately per time-frequency bin.

[0097] According to an embodiment, the signal separator 110 can for example be configured to separate the at least two audio input signals into the first signal portion and the second signal portion smoothly over time.

[0098] In an embodiment, the signal separator 110 can for example be configured to smooth the separation of the at least two audio input signals over time depending on an attack time defining an adaptation of the separation mask when the coherence increases and / or a release time defining an adaptation of the separation mask when the coherence decreases.

[0099] According to an embodiment, the signal separator 110 can for example be configured to employ a different attack time for positively correlated signals than for negatively correlated signals; and / or the signal separator 110 can for example be configured to employ a different release time for positively correlated signals than for negatively correlated signals.

[0100] In an embodiment, the signal separator 110 can for example be configured to smooth a change of the attack time over time; and / or the signal separator 110 can for example be configured to smooth a change of the release time over time.

[0101] According to an embodiment, the signal separator 110 can for example be configured to change the attack time only up to a first predetermined amount within a first predetermined time period; and / or wherein the signal separator 110 can for example be configured to change the attack time only up to a second predetermined amount within a second predetermined time period. The second predetermined amount can for example be equal to or different from the first predetermined amount; and wherein the second predetermined time period can for example be equal to or different from the first predetermined time period.

[0102] In an embodiment, the apparatus can for example be configured to process only a certain frequency band of the at least two audio input signals.

[0103] According to an embodiment, the apparatus can for example be configured to process only a certain signal portion of the at least two audio input signals that exhibits a certain signal characteristic or that exhibits a certain property.

[0104] In an embodiment, the certain signal property or certain attribute of an audio input signal of the at least two audio input signals may, for example, be at least one of:

[0105] the presence of speech,

[0106] the presence of a sound portion,

[0107] whether the audio input signal is a center signal,

[0108] whether the audio input signal is received or derived from other channels as a center signal,

[0109] whether the audio input signal is an ambient signal,

[0110] whether the audio input signal is a channel signal,

[0111] whether the audio input signal is an object signal,

[0112] whether the audio input signal is a surround signal,

[0113] direction information of the audio input signal,

[0114] sound image positioning information of the audio input signal,

[0115] whether the audio input signal contains a transient signal portion.

[0116] According to an embodiment, the signal separator 110 may, for example, be configured to determine the correlation indicators in the time domain.

[0117] In an embodiment, the signal separator 110 may, for example, be configured to calculate frequency-dependent correlation indicators in the time domain by employing a filter bank and by calculating certain frequency band correlations.

[0118] According to an embodiment, the apparatus further comprises a device-specific processing stage for generating a single loudspeaker output from the at least two audio output signals.

[0119] In an embodiment, the apparatus may, for example, be configured to feed the at least two audio output signals into each of three or more loudspeakers.

[0120] According to an embodiment, the apparatus may, for example, be configured to receive information about a loudspeaker setup. The apparatus may, for example, be configured to use the information about the loudspeaker setup to bypass or not bypass the processing by the signal separator 110, the signal processor 120 and the combiner 130.

[0121] In an embodiment, the apparatus further comprises a device-specific processing stage for generating two loudspeaker feeds for the two loudspeakers from the at least two audio output signals using the information about one or more capabilities of the two loudspeakers and / or the information about the distance between the two loudspeakers.

[0122] According to an embodiment, the at least two audio input signals can for example be at least three audio input signals.

[0123] In an embodiment, the apparatus can for example be configured to process the at least three audio input signals by applying the processing of the signal separator 110, the signal processor 120 and the combiner 130 twice or more times.

[0124] According to an embodiment, the apparatus can for example be configured to apply the processing of the signal separator 110, the signal processor 120 and the combiner 130 twice or more times in parallel and / or sequentially.

[0125] In an embodiment, the apparatus can for example be configured to process the at least three audio input signals by extending the processor to multiple inputs and by employing means for selecting two of the three or more audio input signals to be processed according to a selection parameter or a control parameter.

[0126] According to an embodiment, the apparatus can for example be configured to process the at least three audio input signals by computing coherence and correlation between two or more pairs of signals of the at least three audio input signals and / or by computing different combinations thereof; and the signal processor 120 can for example be configured to phase align two or more of the at least three audio input signals using a phase of a first signal portion of one of the at least three audio input signals.

[0127] In an embodiment, the signal separator 110 can for example be configured to separate the at least two audio input signals into the first signal portion and the second signal portion also depending on a gain difference and / or a phase difference between the at least two audio input signals.

[0128] According to an embodiment, the signal separator 110 can for example be configured to separate only a coherent signal portion of the at least two audio input signals having a positive degree of correlation greater than a threshold value into the first signal portion of the at least two audio input signals.

[0129] In an embodiment, the apparatus can for example be configured to process only a coherent signal portion of the at least two audio input signals having a negative degree of correlation less than a threshold value.

[0130] According to an embodiment, the apparatus can for example be configured to smooth the coherence values computed across frequencies.

[0131] In embodiments, the apparatus can be configured to smooth the separation factor along the frequency, for example.

[0132] According to embodiments, the first signal part can be a coherent signal part, for example, and / or the second signal part can be a non-coherent signal part, for example.

[0133] In the following, specific embodiments of the invention will be described.

[0134] The inventive method describes a processor which receives two audio signals at its input.

[0135] The signals are analyzed and processed in order to prevent adverse effects in possible subsequent processing or subsequent processing steps.

[0136] Figure 4 A scenario is shown in which a processor receives two audio signals at its input, processes them and outputs two audio signals.

[0137] In the preferred embodiments described in the following, the signal parts which would cause adverse effects are distinguished from the signal parts which would not cause adverse effects by analyzing the similarity between the two signals. The similarity between the two signals is estimated based on correlation and coherence, as explained further below.

[0138] After the input signals have been analyzed, the phase information of the coherent parts is aligned.

[0139] The phase alignment of the coherent parts comprises adjusting the phase information of one of the two signals so that it matches the phase of the other signal.

[0140] Two variants are possible:

[0141] 1. The phase alignment of the two signals is such that the anti-phase signals or anti-phase signal parts have the same phase after processing.

[0142] 2. The phase alignment of the two signals is such that the in-phase signals or in-phase signal parts have inverted phase after processing.

[0143] Figure 5 Further details of the audio signal processing according to embodiments are shown.

[0144] In the preferred embodiments, successive short parts of the signals are converted into the frequency domain (STFT module in Figure 5 STFT = "Short-Time Fourier Transform").

[0145] The input signals In_1 and In_2 are subjected to a separation process, which decomposes the input signals into a coherent part (Coh_1 and Coh_2) and a non-coherent part (Noncoh_1 and Noncoh_2) of the two input signals.

[0146] The coherent signal part of one of the input signals (Coh_2) is processed such that its signal part is phase-aligned with the phase of Coh_1.

[0147] At this processing stage, the amplitude of Coh_2 is not changed.

[0148] For Coh_1, neither the amplitude nor the phase is changed.

[0149] The non-coherent parts of the two input signals are also left unchanged.

[0150] After phase alignment, the coherent parts of the signals are combined with the non-coherent parts.

[0151] The processed signal (Proc_2 + Noncoh_2) is further processed such that the signal energy (time / frequency dependent, e.g. per frequency bin or band) of the signal Out_2 corresponds to the signal energy of the input signal In_2. The specific band type is irrelevant. It can be, for example, a octave band, a 1 / 3 octave band, a Bark scale band, etc. This applies to all other processing steps performed in the time / frequency domain as well. (This is not necessary for Out_1, since Out_1 corresponds to In_1.)

[0152] Although this power compensation is performed in the preferred embodiment, it is generally optional, since in many use cases the described phase alignment does not result in a significant energy change of the processed signal compared to the original signal.

[0153] The output signal is then converted back to the time domain.

[0154] The individual processing steps in a specific embodiment will be described in the following.

[0155] The signal separation is based on the calculation of a separation mask (M(f,s)) in the time / frequency domain. This separation mask contains values between 0 and 1 in each frequency bin of each time frame. By multiplying the input spectrum with the separation mask, the "coherent" spectrum (Coh_1(f,s), Coh_2(f,s)) and the "non-coherent" spectrum (Noncoh_1(f,s), Noncoh_2(f,s)) can be obtained:

[0156]

[0157]

[0158]

[0159]

[0160] where f is an indication of the discrete frequency (bin) and s an indication of the discrete time (for the time frame).

[0161] The separation mask is computed based on the coherence between the two input signals.

[0162] The coherence between the signals In_1 and In_2 can be computed based on the average self-spectrum and cross-spectrum, where the time averaging process is controlled by the factor a (which determines the influence of past signal behavior on the current estimate).

[0163]

[0164]

[0165]

[0166]

[0167] where indicates the expected value, indicates the complex conjugate.

[0168] The average / expected value of a single frame s is obtained from the value of the frame s and the average of the previous frame:

[0169]

[0170] The coherence value for each time-frequency bin is obtained this way.

[0171] The coherence can take values between 0 and 1, where

[0172] • a value of 0 indicates that the two input signals are uncorrelated, which means that they are independent from each other.

[0173] • a coherence value of 1 indicates a perfect positive or perfect negative correlation. This shows that the signals are either identical (coherence = 1 and correlation = 1) or they carry the same signal, but one of the signals is phase-inverted compared to the other signal (coherence = 1 and correlation = -1).

[0174] (The terms angle and phase can be used interchangeably.)

[0175] From the normalized absolute angle of the cross-spectrum an indicator of the sign of the correlation between the signals In_1 and In_2 can be obtained. It takes values between 0 and 1, where values close to 0 indicate a positive correlation (= phase difference close to zero), values close to 1 indicate a negative correlation (= phase difference close to 180 degrees), and values close to 0.5 indicate no correlation or a phase difference of 90 degrees. To obtain this indicator, the frequency-dependent absolute cross-spectrum phase is summarized to a single value by a frequency-weighted average.

[0176] In Figure 6 In a preferred embodiment of the processing shown, the weighting factor decreases with increasing frequency.

[0177] The correlation indicator (C) can be calculated as follows,

[0178]

[0179] where is the absolute normalized cross-spectrum angle, is a frequency-dependent weighting factor.

[0180] In an embodiment, the signal separation can be adjusted by a parameter to define, for example, above which threshold a signal portion is attributed to the coherent or incoherent portion. The separation can be parameterized such that the separation decision is not a binary decision, but the assignment of the coherent and incoherent portions is made with a smooth transition.

[0181] The separation of the coherent and incoherent signal portions is based on an estimated coherence value for each time-frequency bin.

[0182] In a preferred embodiment, a separation function according to the function shown in Figure 7 is used.

[0183] A threshold value for the coherence can be set, which defines the coherence value when half of the signal amplitude (i.e. the amplitude of a particular bin) is attributed to the coherent portion and the other half to the incoherent portion. Depending on the set threshold value, the shape of the smooth transition region for coherence values close to 0 or 1 also changes. (The threshold value cannot be 0 or 1.)

[0184]

[0185]

[0186] where is a frequency-dependent threshold value, is a factor that adjusts the steepness of the extraction curve. can be any number greater than 0 and less than 1, can be any number greater than 0. (For clarity, the frequency and time indices are omitted in the above formula.) ​

[0187] The separation function can be different in different frequency regions, i.e. it can be adjusted in a frequency dependent manner. (For example, for low frequencies, It can be set, for example, to a lower floor value, and / or, for high frequencies, to a higher floor value, with a linear transition between the lower floor value and the higher floor value, wherein It can be set, for example, to a value between 2 and 15.

[0188] The threshold and the specified separation function can be different in different frequency regions, i.e. they can be adjusted in a frequency dependent manner.

[0189] In a preferred embodiment, frequency dependent thresholds are defined for low frequencies and high frequencies, respectively, with a linear transition between the two.

[0190] For example, a lower threshold can be used in the low frequency range, while a higher threshold can be used in the high frequency range, so that only the highly correlated signal parts will end up in the coherent part of the separated signal, which will eventually be phase-aligned.

[0191] In an embodiment, the (frequency dependent) threshold is a parameter that can be adjusted depending on the specific application scenario.

[0192] The final factor used to separate the signal parts into coherent and incoherent parts can also be smoothed over time, to achieve a smooth transition for different coherence signals and to avoid rapid fluctuations in the separation.

[0193] In a preferred embodiment, the smoothing employs different time constants: an attack time constant when changing from low to high coherence, and a release time when changing from a high to a low coherence region.

[0194] The attack time and the release time are adjustment parameters that control the speed of adaptation of the separation mask. If the signal contains suddenly appearing coherent content, a short attack time will cause the value of the separation mask to increase rapidly between two consecutive time frames; if the coherent signal is no longer coherent, a short release time will cause the value of the separation mask to decrease rapidly between two consecutive frames.

[0195] A long attack time and a long release time cause the separation mask to adapt slowly to changes in the content.

[0196] The positively correlated signal part and the negatively correlated signal part can employ different attack times and release times.

[0197] Different attack and release times are applied for positively or negatively correlated signals in order to adapt the mask adaptively to the actual signal content. In a preferred embodiment, the correlation sign is used to control the attack and release times.

[0198] Figure 8 A graph according to an embodiment is shown, which depicts an example mapping between the aforementioned correlation indicator and a correlation adaptive time constant (which can be an attack time or a release time). For this mapping, a short time constant of 10 ms is defined for negatively correlated content with a correlation indicator higher than 0.8, while a time constant of 300 ms is defined for positively correlated content with a correlation indicator lower than 0.5.

[0199] Since the correlation of a signal can change quickly between subsequent frames, it is desirable to change the attack and release times quickly in response to the correlation indicator. Figure 8 The target attack and release times associated with the correlation indicator shown function can also change quickly. However, it is not desirable to change the attack and release times controlling the speed of adaptation of the separate mask too quickly between subsequent frames.

[0200] Therefore, additional time constants are employed to control the speed of adaptation of the aforementioned attack and release times. They function the same as the aforementioned attack and release times, except that they control the speed of adaptation of the attack and release time values, not the adaptation of the separate mask values.

[0201] Figure 9 An example of smoothing of the attack and release times according to an embodiment is shown. Figure 9 In particular, the behavior of a shorter attack time and a longer release time in the time constant smoothing is shown.

[0202] Since it is not desirable to change the separate attack and release times too quickly, the actual applied values are smoothed in time to avoid sudden changes, and the speed of adaptation of the attack and release times can be controlled by a smoothing parameter.

[0203] After the signals are separated in this way, the phases of the coherent signal parts are aligned.

[0204] One of the signals can be chosen as a reference signal (In_1 is chosen as the reference in our example. Since only the similar signal parts in both signals are processed, the actual choice is not crucial).

[0205] In a preferred embodiment, the alignment comprises:

[0206] • For variant 1 : the phase information of the reference signal is copied to the other signal (i.e. both signals use the phase information of the reference signal).

[0207] • For variant 2: The phase information of the reference signal is copied and inverted to the other signal.

[0208] The non-coherent part of the signal to be processed (and the coherent part not fed into the phase alignment) remains unchanged.

[0209] Since the preferred way of processing is in the time-frequency domain, all processing parameters can be set to be frequency dependent. For example, the time constant in the coherence calculation , the coherence threshold, the signal separation time constant control, the correlation indicator, etc. can be adjusted frequency dependent, depending on the specific application scenario.

[0210] Similarly, the whole processing can be performed only in selected frequency bands.

[0211] An alternative way of limiting the processing to specific parts of the input would be to apply the processing dependent on the signal, i.e. for example only to the speech or human voice parts of the input signal.

[0212] In an alternative embodiment, the correlation indicator can be calculated for example in the time domain (which would correspond to the actual signal correlation). With the filter bank and the correlation calculation per frequency band, a frequency dependent correlation calculation can be achieved in the time domain.

[0213] The (frequency dependent) correlation or correlation indicator can be used to extract only the parts of the content that are within a certain correlation limit, i.e. only the parts of the signal that are within a specified correlation range are separated to be fed to Coh_1 and Coh_2.

[0214] In the following, various application scenarios for the specific embodiments are given.

[0215] In a known processing path (for example in a production or reproduction system), the processing can be applied to signals that are later combined.

[0216] In a system where the processing or reproduction path can change during operation (for example due to user adjustable, adaptive to specific boundary conditions or environment), the processing can be applied to signals that are later most likely combined.

[0217] Figure 10 An exemplary application according to the embodiments is shown (in the block marked PFCP in Figure 10 , where a simple smart speaker as introduced before is used.

[0218] A specific advantage of applying the pre-processing to the input signal (for example compared to applying the pre-processing as a last step of device specific processing) can be embodied by a multi-playback device application scenario.

[0219] For example, the method can be particularly beneficially applied if the target playback system consists of multiple playback devices. This can be the case, for example, in a multi-room playback scenario when multiple smart speakers are fed the same binaural signal. It is sufficient to apply the processing only once (e.g. in the player device, or in a master device that feeds the signal to other devices).

[0220] Figure 11 Such a scenario is shown according to an embodiment, in which an audio signal is sent from a source device (like a media player or a receiver, etc.) to multiple playback devices (like smart speakers).

[0221] These multiple playback devices can be distributed over different rooms, or can be playing in unison in the same room.

[0222] Thus, the pre-processing only needs to be performed once, without having to do it separately in each playback device.

[0223] In a system that is capable of acquiring information about the loudspeaker setup (e.g. the type of loudspeakers, the position of the loudspeakers), this information (LSMD-loudspeaker setup metadata) can be fed back to the processor to guide the actual processing.

[0224] If the actual playback setup is capable of reproducing the originally intended hearing effects when playing back the signal, it is even possible to use this to switch off or bypass the processing.

[0225] Similarly, in a system that is capable of acquiring information about the playback environment (e.g. the room size, the room shape, the room acoustics), this information (LSMD-playback environment metadata) can be fed back to the processor to guide the actual processing. Figure 12 In the depicted exemplary device according to an embodiment (having two loudspeaker drivers in a single enclosure), the frequency selection processing can be adjusted to the specifications and characteristics of this particular device. Depending on the loudspeaker capabilities and the distance between the two loudspeakers, certain hearing effects of the anti-phase signal can be reproduced, while in other frequency ranges they will result in a cancellation of the signal content.

[0226] Since the playback devices are known in advance, the processing parameters can be tuned or adjusted accordingly.

[0227] (Such adaptation can be achieved by manual adjustment, or based on parameters of the playback system, e.g. reproducible frequency range, distance between loudspeakers, etc.)

[0228] Figure 13 An exemplary application of the described processing (PFCP module) in the previously introduced soundbar device according to an embodiment is shown.

[0229] Using the method for multiple input signals (e.g. more than two) can be done by:

[0230] • By applying the processing multiple times, it can be done in a parallel fashion (see Figure 14), in series / sequential manner (see Figure 15 ) or a mix of both (see Figure 16 );

[0231] • By extending the processor for multiple inputs and adding means for selecting which two input channels should be processed based on selection parameters or control parameters (see Figure 17 );

[0232] • Calculating the coherence and correlation between multiple channels and different combinations thereof and modifying the processor such that phase alignment from one channel to multiple other channels is performed (see Figure 18 ).

[0233] The method can advantageously be used in many applications.

[0234] In embodiments, the application of the power compensation can be decided based on considerations of criteria or parameters regarding the design, complexity, performance or quality of the target system. Figure 19 A processing without power compensation is shown.

[0235] The separation mask is calculated from the coherence between the two input signals. In alternative embodiments of the processing, the separation mask can take into account additional information, such as the gain difference and the phase difference between the two channels. In a preferred embodiment of the processing, only the parts of the signals that need to be phase adjusted to avoid adverse effects in subsequent processing steps are extracted.

[0236] In some embodiments, it can be advantageous to apply the signal separation only in a limited frequency range, meaning that all signal parts outside the specified frequency range will eventually be attributed to the incoherent part and not be subject to further processing.

[0237] In alternative embodiments of the processing, only coherent signal parts with a high positive (or negative) correlation are processed. In another alternative embodiment of the processing, only coherent signal parts with a high potential for cancellation upon (phase-correct or phase-inverted) summation are processed.

[0238] In alternative embodiments where only parts of the coherent signal parts are phase aligned, both of the aforementioned phase alignment variants are equally feasible.

[0239] The coherence values calculated in the STFT domain can also be cross- frequency smoothed.

[0240] In alternative embodiments, the separation / extraction factor can also be smoothed along the frequency.

[0241] According to other embodiments, it is not always desirable to feed all the coherent content into the first signal part. It is only necessary to phase adjust the parts of the signal that can have adverse effects in subsequent processing.

[0242] For example, in the case of processing a signal that is ultimately reproduced through a single loudspeaker, part of these related signals are the anti-phase parts.

[0243] The part of the (coherent) signal that does not cancel out upon summation can be removed from the separation mask. The cancellation is calculated as follows, where and are the average auto-spectra of the input signals and is the average auto-spectrum of the complex sum of X and Y.

[0244]

[0245]

[0246]

[0247]

[0248] The calculated cancellation per frequency bin in each frame (in dB) is converted to a factor which is applied to the separation mask to obtain the modified separation mask . For small amounts of cancellation, the separation mask values are decreased, i.e. less signals are phase-aligned.

[0249] The conversion function F can take the shape as shown in Figure 20 where signals that result in a cancellation above a certain threshold are not removed from the separation mask, while signals that result in a cancellation below that threshold are gradually removed more from the separation mask.

[0250]

[0251]

[0252] Although some aspects are described in the context of an apparatus, it is clear that these aspects also correspond to a description of the corresponding method, where a module or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step, also represent a description of a corresponding module, means, or feature of corresponding apparatuses. Some or all of the method steps can be executed by (or using) hardware apparatus, like for example, microprocessors, programmable computer or electronic circuitry. In some embodiments, one or more of the most important method steps can be executed by such an apparatus.

[0253] ​Depending on certain implementation requirements of the inventive methods, embodiments of the application can be implemented in hardware or in software, or in a combination of hardware and software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which can cooperate with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.

[0254] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein can be performed.

[0255] Generally, embodiments of the present application can be implemented as a computer program product, comprising a computer readable medium having stored thereon the program code for performing one of the methods described herein. The program code can be stored in a machine readable carrier.

[0256] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a carrier.

[0257] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program is executed on a computer.

[0258] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.

[0259] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0260] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein.

[0261] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0262] Other embodiments of the invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0263] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0264] The apparatus described herein can be implemented using hardware devices, computers, or a combination of hardware devices and computers.

[0265] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0266] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations in the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent is limited only by the scope of the appended claims and not by the specific details presented in the description and explanation of the embodiments herein.

Claims

1. An apparatus for audio signal processing, wherein the apparatus comprises: A signal separator (110) is used to separate each of at least two audio input signals into a first signal portion and a second signal portion; A signal processor (120) is configured to obtain a phase-aligned signal portion of each of the at least two audio input signals from a first signal portion of each of the at least two audio input signals by modifying a first signal portion of at least one of the at least two audio input signals; The signal processor (120) is configured to modify the first signal portion of the at least one audio input signal by phasing the first signal portion of the at least one audio input signal with the first signal portion of at least one of the at least two audio input signals; and A combiner (130) is used to combine the phase-aligned signal portion and the second signal portion of each of the at least two audio input signals to obtain at least two audio output signals.

2. The apparatus according to claim 1, The signal separator (110) is configured to separate each of the at least two audio input signals into a first signal portion and a second signal portion based on coherence and / or correlation.

3. The apparatus according to claim 1, The signal separator (110) is configured to separate each of the at least two audio input signals into a first signal portion and a second signal portion based on the coherence and / or correlation between the first signal portion and the signal portions of one or more other audio input signals among the at least two audio input signals.

4. The apparatus according to any one of the preceding claims, The signal splitter (110) is configured to split each of at least two audio input signals into a first signal portion and a second signal portion, such that the first signal portion is coherent with the signal portions of one or more other audio input signals among the at least two audio input signals.

5. The apparatus according to any one of the preceding claims, in, In order to obtain the phase-aligned signal portion of each of the at least two audio input signals, the signal processor (120) is configured to modify a first signal portion of the at least one audio input signal and is configured not to modify the first signal portion of the at least one other audio input signal.

6. The apparatus according to any one of claims 1 to 4, in, In order to obtain the phase alignment signal portion of each of the at least two audio input signals, the signal processor (120) is configured to modify a first signal portion of each of the at least two audio input signals.

7. The apparatus according to any one of the preceding claims, The at least two audio input signals are exactly two audio input signals. The at least one audio input signal is exactly an audio input signal. The at least one other audio input signal is exactly one other audio input signal, and The at least two audio output signals are exactly two audio output signals.

8. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align a first signal portion of the at least one audio input signal with a first signal portion of the at least one other audio input signal in the frequency domain.

9. The apparatus according to claim 8, The signal processor (120) is configured to align the phase of at least one frequency band of a first signal portion of the at least one audio input signal with the phase of at least one frequency band of a first signal portion of the at least one other audio input signal in the frequency domain.

10. The apparatus according to claim 8, The signal processor (120) is configured to align the phase of each of two or more frequency bands of the first signal portion of the at least one audio input signal with the phase of each of the two or more other frequency bands of the first signal portion of the at least one other audio input signal in the frequency domain.

11. The apparatus according to any one of claims 8 to 10, The device further includes a time-frequency conversion unit for converting the at least two audio input signals represented in the time domain from the time domain to the frequency domain, and The device further includes a frequency-time conversion unit for converting the at least two audio output signals represented in the frequency domain from the frequency domain to the time domain.

12. The apparatus according to claim 11, The time-frequency conversion unit is configured to perform a short-time Fourier transform to transform the at least two audio input signals from the time domain to the frequency domain, and The frequency-time conversion unit is configured to perform a short-time inverse Fourier transform to convert the at least two audio output signals from the frequency domain to the time domain.

13. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align a first signal portion of the at least one audio input signal with a first signal portion of the at least one other audio input signal, such that when the first signal portion of the at least one audio input signal is negatively correlated with the first signal portion of the at least one other audio input signal, the phase-aligned signal portions of the at least one audio input signal and the phase-aligned signal portions of the at least one other audio input signal have the same phase after phase alignment.

14. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align a first signal portion of the at least one audio input signal with a first signal portion of the at least one other audio input signal, such that, when the first signal portion of the at least one audio input signal is positively correlated with the first signal portion of the at least one other audio input signal, the phase-aligned signal portions of the at least one audio input signal and the phase-aligned signal portions of the at least one other audio input signal have inverted phases after phase alignment.

15. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal by copying the phase information of the first signal portion of the at least one other audio input signal to the first signal portion of the at least one audio input signal.

16. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align the first signal portion of the at least one audio input signal with the first signal portion of the at least one other audio input signal by copying and phase-inverting the phase information of the first signal portion of the at least one other audio input signal to the first signal portion of the at least one audio input signal.

17. The apparatus according to any one of the preceding claims, The signal processor (120) is configured to phase-align a first signal portion of the at least one audio input signal with a first signal portion of the at least one other audio input signal without changing the amplitude of the first signal portion of the at least one audio input signal or the amplitude of the first signal portion of the at least one other audio input signal.

18. The apparatus according to any one of the preceding claims, The second signal portion of each of the at least two audio input signals is not modified when combined by the combiner (130).

19. The apparatus according to any one of the preceding claims, The at least two audio input signals include one or more audio channel signals and / or one or more audio object signals and / or one or more surround stereo signals.

20. The apparatus according to any one of the preceding claims, The device includes a power compensator. The total signal energy of the at least two audio output signals corresponds to the total signal energy of the at least two audio input signals, or The signal energy of one of the at least two audio output signals corresponds to the signal energy of one of the at least two audio input signals, or The signal energy of each of the at least two audio output signals corresponds to the signal energy of one of the at least two audio input signals.

21. The apparatus according to claim 20, The power compensator is configured to perform power compensation by frequency range or by frequency band.

22. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to separate each of at least two audio input signals into a first signal portion and a second signal portion by applying a first mask value to the time-frequency interval of the audio input signal to obtain the time-frequency interval of the first signal portion, and by applying a second mask value depending on the first mask value to the time-frequency interval of the audio input signal to obtain the time-frequency interval of the second signal portion.

23. The apparatus according to claim 22, The signal separator (110) is configured to apply the same first mask value to two or more time-frequency intervals of the same frequency band of the audio input signal to obtain two or more time-frequency intervals of a first signal portion of the same frequency band; and / or The signal splitter (110) is configured to apply the same second mask value to two or more time-frequency intervals of the same frequency band of the audio input signal to obtain two or more time-frequency intervals of the second signal portion of the same frequency band.

24. The apparatus according to claim 22 or 23, The signal separator (110) is configured to separate each of at least two audio input signals into a first signal portion and a second signal portion by multiplying the first mask value by the time-frequency interval of the audio input signal to obtain the time-frequency interval of the first signal portion, and by multiplying the second mask value by the time-frequency interval of the audio input signal to obtain the time-frequency interval of the second signal portion, wherein the first mask value presents a value v1, where 0 ≤ v1 ≤ 1, and wherein the second mask value v2 = 1 - v1.

25. The apparatus according to any one of claims 22 to 24, The signal separator (110) is configured to separate a coherent signal portion into a first signal portion and a second signal portion of the audio input signal, such that the first signal portion includes only the coherent signal portions of the at least two audio input signals whose sum is greater than a threshold and which may cancel each other out (e.g., phase correct or phase reversed).

26. The apparatus according to claim 25, The signal separator (110) is configured to update the first mask value and the second mask value for each of a plurality of time-frequency intervals of the audio input signal, such that the first signal portion includes only the coherent signal portions of the at least two audio input signals whose sum is greater than a threshold and which may cancel each other out.

27. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to separate each of at least two audio input signals into a first signal portion and a second signal portion based on the coherence of each of a plurality of time-frequency intervals, wherein the coherence is time-averaged.

28. The apparatus according to claim 27, The signal separator (110) is configured to determine the coherence of each of the plurality of time-frequency intervals based on the time-averaged autocorrelation of the time-frequency intervals and the time-averaged crosscorrelation of the time-frequency intervals.

29. The apparatus according to claim 27 or 28, The signal separator (110) is configured to determine a frequency-dependent absolute cross-spectral phase, which is summed into a single absolute cross-spectral phase value by a frequency-dependent mean.

30. The apparatus according to claim 29, The absolute cross-spectral phase value of a single frequency correlation, which is 0, indicates a positive correlation. The absolute cross-spectral phase value with a single frequency correlation of 0.5 indicates no correlation, and The absolute cross-spectral phase value with a value of 1 for a single frequency correlation indicates a negative correlation.

31. The apparatus according to any one of claims 27 to 30, The signal separator (110) is configured to separate each of at least two audio input signals into a first signal portion and a second signal portion by employing a separation function, the separation function depending on the coherence of the time-frequency intervals in a plurality of time-frequency intervals.

32. The apparatus according to claim 31, The separation function described therein separates the amplitude of the time-frequency interval into a coherent amplitude component and an incoherent amplitude component.

33. The apparatus according to claim 31 or 32, The separation function mentioned above is frequency-dependent.

34. The apparatus according to any one of claims 31 to 33, The separation function depends on the signal properties of at least one of the at least two audio input signals.

35. The apparatus according to any one of claims 31 to 34, The separation function mentioned above depends on the threshold.

36. The apparatus according to claim 35, The threshold is frequency-dependent, such that the signal separator (110) is configured to, for the same coherence, allocate a larger amplitude portion to the first signal portion presenting a lower frequency compared to the amplitude portion allocated to the first signal portion presenting a higher frequency.

37. The apparatus according to claim 35 or 36, The device includes an interface for setting the threshold.

38. The apparatus according to claim 37, The interface is configured to set the threshold individually by frequency band or individually by time-frequency interval.

39. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to smooth the separation of the at least two audio input signals into a first signal portion and a second signal portion over time.

40. The apparatus according to claim 39, The signal separator (110) is configured to smooth the separation of the at least two audio input signals over time according to the on-time and / or release time, wherein the on-time defines the adaptation of the separation mask as coherence increases and the release time defines the adaptation of the separation mask as coherence decreases.

41. The apparatus according to claim 40, The signal separator (110) is configured to use different on-times for positively correlated signals and for negatively correlated signals; and / or The signal separator (110) is configured to use different release times for positively correlated signals and for negatively correlated signals.

42. The apparatus according to claim 41, The signal splitter (110) is configured to smooth the change in the onset time over time; and / or The signal separator (110) is configured to smooth the change of the release time over time.

43. The apparatus according to claim 41 or 42, The signal splitter (110) is configured to change the on-time only by a maximum of a first predetermined amount during a first predetermined time period; and / or the signal splitter (110) is configured to change the on-time only by a maximum of a second predetermined amount during a second predetermined time period; Wherein the second predetermined quantity is equal to or different from the first predetermined quantity; and wherein the second predetermined time period is equal to or different from the first predetermined time period.

44. The apparatus according to any one of the preceding claims, The device is configured to process only a specific frequency band of the at least two audio input signals.

45. The apparatus according to any one of the preceding claims, The device is configured to process only a specific portion of the at least two audio input signals that exhibits specific signal characteristics or properties.

46. ​​The apparatus according to claim 45, The specific signal characteristic or specific property of the audio input signal in the at least two audio input signals is at least one of the following: There is voice recording. There is a sound component. Is the audio input signal a center signal? Whether the audio input signal, which serves as the central signal, is received or derived from other channels. Is the audio input signal an ambient signal? Is the audio input signal a channel signal? Is the audio input signal an object signal? Is the audio input signal a surround sound signal? The direction information of the audio input signal The acoustic image localization information of the audio input signal. Does the audio input signal include a transient signal component? 47. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to determine a correlation indicator in the time domain.

48. The apparatus according to claim 47, The signal separator (110) is configured to calculate a frequency-dependent correlation indication in the time domain by employing a filter bank and by performing a specific frequency band correlation calculation.

49. The apparatus according to any one of the preceding claims, The device further includes a device-specific processing stage for generating a single speaker output from the at least two audio output signals.

50. The apparatus according to any one of claims 1 to 48, The device is configured to feed the at least two audio output signals into each of three or more speakers.

51. The apparatus according to any one of claims 1 to 48, The device is configured to receive information about speaker settings. The device is configured to use the information about the speaker settings to bypass or not bypass the processing performed by the signal splitter (110), the signal processor (120), and the combiner (130).

52. The apparatus according to any one of claims 1 to 48, The device is configured to receive information about speaker settings. The device is configured to use the information about the speaker settings to bypass or not bypass processing in one or more frequency bands performed by the signal splitter (110), the signal processor (120), and the combiner (130).

53. The apparatus according to any one of claims 1 to 48, The device further includes a device-specific processing stage for generating two speaker feeds for the two speakers from the at least two audio output signals, using information about one or more capabilities of the two speakers and / or information about the distance between the two speakers.

54. The apparatus according to any one of the preceding claims, The at least two audio input signals are at least three audio input signals.

55. The apparatus according to claim 54, The device is configured to process the at least three audio input signals by applying the processing of the signal splitter (110), the signal processor (120), and the combiner (130) two or more times.

56. The apparatus according to claim 55, The device is configured to apply the processing of the signal splitter (110), the signal processor (120), and the combiner (130) in parallel and / or sequentially two or more times.

57. The apparatus according to claim 54, The device is configured to process the at least three audio input signals by extending the processor to multiple inputs and by employing means for selecting two audio input signals to be processed from three or more audio input signals according to selection parameters or control parameters.

58. The apparatus according to claim 54, The device is configured to process the at least three audio input signals by calculating the coherence and correlation between two or more pairs of signals among the at least three signals, and / or by calculating different combinations thereof; and the signal processor (120) is configured to phase-align two or more other audio input signals among the at least three audio input signals by using the phase of a first signal portion of one of the audio input signals.

59. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to further separate the at least two audio input signals into a first signal portion and a second signal portion based on the gain difference and / or phase difference between the at least two signals.

60. The apparatus according to any one of the preceding claims, The signal separator (110) is configured to separate only the coherent signal portions of at least two audio input signals that have a positive correlation greater than a threshold into a first signal portion of at least two audio input signals.

61. The apparatus according to any one of the preceding claims, The device is configured to process only the coherent signal portions of at least two audio input signals whose negative correlation is less than a threshold.

62. The apparatus according to any one of the preceding claims, The device is configured to smooth the coherence values ​​calculated across frequencies.

63. The apparatus according to any one of the preceding claims, The device is configured to smooth the separation factor along the frequency.

64. The apparatus according to any one of the preceding claims, Wherein the first signal portion is a coherent signal portion, and / or wherein the second signal portion is an incoherent signal portion.

65. A method for audio signal processing, wherein the method comprises: Separate each of at least two audio input signals into a first signal portion and a second signal portion; By modifying the first signal portion of at least one of the at least two audio input signals, the phase alignment signal portion of each of the at least two audio input signals is obtained from the first signal portion of each of the at least two audio input signals; The modification of the first signal portion of the at least one audio input signal is performed by aligning the first signal portion of the at least one audio input signal with the first signal portion of at least one of the at least two audio input signals. and The phase-aligned signal portion and the second signal portion of each of the at least two audio input signals are combined to obtain at least two audio output signals.

66. A computer program, when executed on a computer or signal processor, for implementing the method of claim 65.