source separation
By processing multi-channel input signals using fractional delay filters and random direction optimization techniques, the problem of audio source separation in microphone received signals is solved, achieving high-precision audio signal separation.
Patent Information
- Application Number
- CN202080070200.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-04
- Filing Date
- 2020-10-02
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2040-10-02
AI Technical Summary
Existing technologies struggle to effectively separate signals from different audio sources in multi-channel input signals, especially when crosstalk exists in the signals received by the microphone, making it difficult to accurately extract their respective audio signals.
By using fractional delay filters and random direction optimization techniques, combined with signal delay and scaling processing, multi-channel input signals are separated to obtain their respective output signals. The signal processing process is optimized using formulas (6) and (8) to reduce crosstalk.
It achieves effective separation of different audio sources in multi-channel input signals, reduces signal crosstalk, and improves the separation accuracy and quality of audio signals.
Smart Images

Figure GDA0005645967340000081 
Figure GDA0005645967340000082 
Figure GDA0005645967340000091
Abstract
Description
Technical Field
[0001] This example relates to methods and apparatus for obtaining multiple output signals associated with different sources (e.g., audio sources). This example also relates to methods and apparatus for signal separation. This example further relates to methods and apparatus for teleconferencing. Techniques for separation (e.g., audio source separation) are also disclosed. Techniques for fast time-domain stereo audio source separation (e.g., using fractional delay filters) are also discussed. Background Technology
[0002] Figure 1 The setup of the microphone, indicated by 50a, is shown. Microphone 50a may include a first microphone mic0 and a second microphone mic1, which are shown here as being 5cm (50mm) apart from each other. Other distances are also possible. Two different sources (source0 and source1) are shown here. As indicated by angles β0 and β1, they are placed in different positions (and also in different orientations relative to each other).
[0003] Multiple input signals M0 and M1 (from the microphone, also collectively indicated as multi-channel or stereo input signal 502) are obtained from sound sources source0 and source1. Source0 generates the audio sound with index S0, while source1 generates the audio sound with index S1.
[0004] Microphone signals M0 and M1 can be considered, for example, as input signals. Multichannel with more than two channels can be considered instead of stereo signal 502.
[0005] In some real-world examples, there can be more than two input signals (e.g., additional input channels besides M0 and M1), although this primarily focuses on two channels. Nevertheless, this example works for any multi-channel input signal. In this example, signals M0 and M1 are not necessarily obtained directly from a microphone, as they can be obtained, for example, from a stored audio file.
[0006] Figure 2a and Figure 4 The diagram illustrates the interaction between sources source0 and source1 and microphones mic0 and mic1. For example, source0 generates audio sound S0, which primarily reaches microphone mic0 and also microphone mic1. The same applies to source1, whose generated audio source S1 primarily reaches microphone mic1 and also microphone mic0. Figure 2a and Figure 4As can be seen, the time required for sound S0 to reach microphone mic0 is less than the time required to reach microphone mic1. Similarly, the time required for sound S1 to reach mic1 is less than the time required to reach mic0. The strength of signal S0 is generally attenuated when it reaches microphone mic1 compared to when it reaches mic0, and vice versa.
[0007] Accordingly, in the multi-channel input signal 502, the channel signals M0 and M1 make the signals S0 and S1 from source0 and source1 a combination of each other. Therefore, a separation technique is sought. Summary of the Invention
[0008] In the following text, the text within square brackets and parentheses indicates non-restrictive examples.
[0009] According to one aspect, an apparatus (e.g., a multi-channel or stereo audio source separation device) is provided for obtaining multiple output signals (S'0, S'1) associated with different sound sources (source0, source1) based on multiple input signals (e.g., microphone signals) that combine signals from different sound sources (source0, source1).
[0010] The device is configured to combine a first input signal [M0] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a1·z] of a second input signal [e.g., M1]. -d1 Combining [·M1] [for example, by subtracting a delayed and scaled version of the second input signal from the first input signal, for example by S'0=M0(z)-a1·z] -d1 ·M1(z)], to obtain the first output signal [S'0];
[0011] The device is configured to compare a second input signal [M1] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a0·z] of a first input signal [M0]. -d0 Combining [·M0] [for example, by subtracting a delayed and scaled version of the first input signal from the second input signal, for example by S'1=M1(z)-a0·z] -d0 ·M0(z)], to obtain the second output signal [S'1];
[0012] The device is configured to determine the objective function using stochastic direction optimization [e.g., by performing one of the operations defined in other embodiments; and / or by finding delay and decay values that minimize the objective function, which may be, for example, the objective function in equations (6) and / or (8)].
[0013] A first scaling value [a0] is used to obtain a delayed and scaled version [a0·z] of the first input signal [M0]. -d0 ·M0];
[0014] A first delay value [d0] is used to obtain a delayed and scaled version [a0·z] of the first input signal [M0]. -d0 ·M0];
[0015] The second scaling value [a1] is used to obtain a delayed and scaled version [a1·z] of the second input signal [M1]. -d1 ·M1];
[0016] as well as
[0017] The second delay value [d1] is used to obtain the delayed and scaled version [a1·z] of the second input signal. -d1 ·M1).
[0018] The delayed and scaled version of the second input signal [M1] [a1·z] -d1 ·M1] can be combined with the first input signal [M0] and obtained by applying a fractional delay to the second input signal [M1].
[0019] The delayed and scaled version of the first input signal [M0] [a0·z] -d0 ·M0] can be combined with the second input signal [M1] and obtained by applying a fractional delay to the first input signal [M0].
[0020] This device can sum multiple products between the following two [for example, as in formula (6) or (8)]:
[0021] -The corresponding element in the first set of normalized amplitude values [P] i (n), where i is 0 or 1 [for example, as in formula (7)].
[0022] [China], and
[0023] - The logarithm of the quotient formed based on the following:
[0024] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0025] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0026] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0027] Random direction optimization can make the candidate parameters form a candidate vector [e.g., with four elements (entries), such as corresponding to a0, a1, d0, d1], which is iteratively refined by modifying the vector in random directions [e.g., in different iterations].
[0028] Random direction optimization can make the candidate parameters form a candidate vector [e.g., with four elements, such as corresponding to a0, a1, d0, d1], which is iteratively refined by modifying the vector in random directions [e.g., in different iterations, see below].
[0029] Random orientation optimization allows a metric and / or value indicating the similarity (or dissimilarity) between the first and second output signals to be measured, and the first and second output measurements are selected as those measurements associated with candidate parameters that indicate the lowest similarity (or highest dissimilarity).
[0030] At least one of the first scaling value and the second scaling value, as well as the first delay value and the second delay value, is obtained by minimizing the mutual information or correlation measurement of the output signal.
[0031] According to one aspect, an apparatus is provided for obtaining multiple output signals [S'0, S'1] associated with different sound sources [source1, source2] based on multiple input signals [e.g., microphone signals] [M0, M1] which combine signals from different sound sources (source0, source1).
[0032] The device is configured to combine a first input signal [M0] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a1·z] of a second input signal [M1]. -d1 The device is configured to combine [M1] to obtain a first output signal [S'0], wherein the device is configured to apply a fractional delay [d1] to the second input signal [M1], wherein the fractional delay (d1) can indicate the signal (H) arriving at the first microphone (mic0) from the second source (source1). 1,0 The delay of ·S1) (e.g., caused by H)1,0 The delay (represented by the signal) and the signal (H) from the second source (source1) to the second microphone (mic1) 1,1 The delay of ·S1) (e.g., caused by H) 1,1 The relationship and / or difference between the delays (represented by the fractional delay d1) and the approximate fraction H. 1,0 (z) / H 1,1 [The exponent of the z-term of the result of (z)];
[0033] The device is configured to compare a second input signal [M1] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a0·z] of a first input signal [M0]. -d0 The first input signal [M0] is combined to obtain a second output signal [S'1], wherein the device is configured to apply a fractional delay [d0] to the first input signal [M0], wherein the fractional delay (d0) can indicate the signal (H) arriving at the first microphone (mic0) from the first source (source0). 0,0 The delay of ·S0) (e.g., caused by H) 0,0 The delay (represented by the signal) and the signal (H) from the first source (source0) to the second microphone (mic1) 0,1 The delay of ·S0) (e.g., caused by H) 0,1 The relationship and / or difference between the delays (represented by the fractional delay) [In the example, the fractional delay d0 can be understood as the approximate fraction H]. 0,1 (z) / H 0,0 [The exponent of the z-term of the result of (z)];
[0034] The device is configured to use optimized determination:
[0035] A first scaling value [a0] is used to obtain a delayed and scaled version [a0·z] of the first input signal [M0]. -d0 ·M0];
[0036] A first fractional delay value [d0] is used to obtain a delayed and scaled version of the first input signal [M0].
[0037] [a0·z -d0 ·M0];
[0038] The second scaling value [a1] is used to obtain a delayed and scaled version [a1·z] of the second input signal [M1]. -d1 ·M1];
[0039] as well as
[0040] The second fractional delay value [d1] is used to obtain a delayed and scaled version of the second input signal [M1].
[0041] [a1·z -d1 ·M1).
[0042] This optimization can be a random direction optimization.
[0043] This device can sum multiple products between the following two [for example, as in formula (6) or (8)]:
[0044] -The corresponding element in the first set of normalized amplitude values [P] i (n), where i is 0 or 1 [for example, as in formula (7)].
[0045] [China], and
[0046] - The logarithm of the quotient formed based on the following:
[0047] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0048] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0049] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0050] According to one aspect, there is provided an apparatus (e.g., a multi-channel or stereo audio source separation device) for obtaining, based on a plurality of input signals (e.g., microphone signals) [M0, M1] from different sound sources [source0, source1] [e.g., by subtracting a delayed and scaled version of a second input signal from a first input signal, and / or by subtracting a delayed and scaled version of a first input signal from a second input signal], multiple output signals [S'0, S'1] associated with said different sound sources [source0, source1].
[0051] The device is configured to combine a first input signal [M0] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a1·z] of a second input signal [M1]. -d1 The first output signal [S'0] is obtained by combining [M1] [for example, by subtracting a delayed and scaled version of the second input signal from the first input signal],
[0052] The device is configured to compare a second input signal [M1] or a processed (e.g., delayed and / or scaled) version thereof with a delayed and scaled version [a0·z] of a first input signal [M0]. -d0 The second output signal [S'1] is obtained by combining [M0] [for example, by subtracting a delayed and scaled version of the first input signal from the second input signal],
[0053] The device is configured to sum multiple products between the following two [e.g., as in formula (6) or (8)]:
[0054] -The corresponding element in the first set of normalized amplitude values [P] i (n), where i is 0 or 1 [for example, as in formula (7)], and
[0055] - The logarithm of the quotient formed based on the following:
[0056] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0057] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0058] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0059] This device can be determined using optimization [e.g., based on "modified KLD calculations"]:
[0060] A first scaling value [a1] is used to obtain a delayed and scaled version of the first input signal [M0].
[0061] A first delay value [d0] is used to obtain a delayed and scaled version of the first input signal.
[0062] The second scaling value [a1] is used to obtain a delayed and scaled version of the second input signal, and
[0063] The second delay value [d1] is used to obtain a delayed and scaled version of the second input signal.
[0064] The first delay value [d0] can be a fractional delay. The second delay value [d1] is a fractional delay.
[0065] This optimization can be a random direction optimization.
[0066] The device can perform at least some of the processes in the time domain. The device can perform at least some of the processes in the z-transform or frequency domain.
[0067] The device can be configured as follows:
[0068] The first input signal [M0] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a1·z] of the second input signal [M1]. -d1 ·M1] combined in the time domain and / or in the z-transform or frequency domain;
[0069] The second input signal [M1] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a0·z] of the first input signal [M0]. -d0 ·M0] is combined in the time domain and / or in the z-transform or frequency domain.
[0070] This optimization can be performed in the time domain and / or in the z-transform or frequency domain.
[0071] The fractional delay (d0) applied to the second input signal [M1] can indicate the relationship and / or difference or arrival between the following two:
[0072] The signal [S0·H] received by the first microphone [mic0] from the first source [source0] 0,0 (z)]; and
[0073] The signal [S0·H] received by the second microphone [mic1] from the first source [source0] 0,1 (z)].
[0074] The fractional delay (d1) applied to the first input signal [M0] can indicate the relationship and / or difference or arrival of the following two:
[0075] The signal [S1·H] received by the second microphone [mic1] from the second source [source1] 1,1 (z)]; and
[0076] The signal [S1·H] received by the first microphone [mic0] from the second source [source1] 1,0 (z)].
[0077] The device can perform optimization [e.g., the optimization causes different candidate parameters [a0, a1, d0, d1] to be iteratively selected and processed, and measures a metric [e.g., as in Equation (6) or (8)] [e.g., on the basis of "modified KLD calculation"] [e.g., objective function] for each candidate parameter, where the metric is a similarity metric (or dissimilarity metric) in order to select a first input signal [M0] and a second input signal [M0] obtained by using candidate parameters [a0, a1, d0, d1] associated with a metric indicating the lowest similarity (or the highest dissimilarity). [Similarity can be conceived as the statistical dependence between the first input signal and the second input signal (or a value associated therewith, such as the value in Equation (7)), and / or dissimilarity can be conceived as the statistical independence between the first input signal and the second input signal (or a value associated therewith, such as the value in Equation (7)).
[0078] For each iteration, the candidate parameters may include a candidate delay (d0) [e.g., a candidate fractional delay] to be applied to the second input signal [M1], which is associated with a candidate relationship and / or candidate difference or arrival between the following:
[0079] The signal [S0·H] received by the first microphone [mic0] from the first source [source0] 0,0 (z)]; and
[0080] The signal [S0·H] received by the second microphone [mic1] from the first source [source0] 0,1 (z)].
[0081] For each iteration, the candidate parameters include a candidate delay (d1) [e.g., a candidate fractional delay] to be applied to the first input signal [M0], which is associated with a candidate relationship and / or candidate difference or arrival between the following two:
[0082] The signal [S1·H] received by the second microphone [mic1] from the second source [source1] 1,1 (z)]; and
[0083] The signal [S1·H] received by the first microphone [mic0] from the second source [source1] 1,0 (z)].
[0084] For each iteration, the candidate parameters may include a candidate relative attenuation value [a0] to be applied to the second input signal [M1], which indicates the candidate relationship and / or candidate difference between the following two:
[0085] The signal [S0·H] received by the first microphone [mic0] from the first source [source0] 0,0 The amplitude of (z)]; and
[0086] The signal [S0·H] received by the second microphone [mic1] from the first source [source0] 0,1 The amplitude of (z)].
[0087] For each iteration, the candidate parameters may include a candidate relative attenuation value [a1] to be applied to the first input signal [M0], which indicates a candidate relationship and / or candidate difference between the following two:
[0088] The signal [S1·H] received by the second microphone [mic1] from the second source [source1] 1,1 The amplitude of (z)]; and
[0089] The signal [S1·H] received by the first microphone [mic0] from the second source [source1] 1,0 The amplitude of (z)].
[0090] The device can change at least one candidate parameter for different iterations by randomly selecting at least one step size [e.g., stochastic direction optimization] from at least one candidate parameter for a previous iteration to at least one candidate parameter for a subsequent iteration.
[0091] The device can randomly select at least one step size [e.g., coeffvariation in line 10 of Algorithm 1] [e.g., random direction optimization].
[0092] This at least one step size can be weighted by pre-selected weights [e.g., coeffweights in line 5 of Algorithm 1].
[0093] This at least one step size is limited by pre-selected weights [e.g., coeffweights in line 5 of Algorithm 1].
[0094] This device allows candidate parameters [a0, a1, d0, d1] to form a candidate vector, wherein, for each iteration, the candidate vector is perturbed [e.g., randomly] by applying a vector of uniformly distributed random numbers [e.g., each between -0.5 and +0.5], which are then element-wise multiplied or added to the candidate vector. [Gradient processing can be avoided][e.g., random direction optimization].
[0095] For each iteration, the candidate vectors are modified (e.g., perturbed) for the step size [e.g., each step size is between -0.5 and +0.5].
[0096] This device allows the number of iterations to be limited to a predetermined maximum number, which is between 10 and 30 (e.g., 20, as in the last three lines of section 2.3).
[0097] The metric can be treated as the Kullback-Leibler divergence.
[0098] Metrics can be based on:
[0099] For each of the first and second signals [M0, M1], the corresponding element [P] in the first set of normalized amplitude values... i (n), where i is 0 or 1 [e.g., as in Equation (7)]. [A trick could be to treat the normalized amplitude values of the time-domain samples as a probability distribution and then measure the metric (e.g., as the Kullback-Leibler divergence, e.g., as obtained by Equation (6) or (8)).]
[0100] For at least one of the first input signal and the second input signal [M0, M1], the corresponding element [P] i [n] can be based on the candidate first or second output signal [S'0, S'1] obtained from the candidate parameters [e.g., as in Equation (7)].
[0101] For at least one of the first input signal and the second input signal [M0, M1], the corresponding element [P] i [n] can be based on the candidate first or second output signal [S'0, S'1] obtained from the candidate parameters [e.g., as in Equation (7)].
[0102] For at least one of the first input signal and the second input signal [M0, M1], the corresponding element [P] i [n] can be obtained as a fraction between the following two:
[0103] The value associated with the candidate first or second output signal [S'0(n), S'1(n)] [e.g., in absolute value form] [e.g., absolute value]; and
[0104] The norm [e.g., 1-norm] associated with the value of the previously obtained first or second output signal [S'0(...n-1), S'1(...n-1)].
[0105] For at least one of the first input signal and the second input signal [M0, M1], the corresponding element [P] i [n] can be obtained by the following formula:
[0106]
[0107] (Here, “s'i(n)” and “s'i” are not written in uppercase because they are not z-transforms in this case.)
[0108] The measure can include the logarithm of the quotient formed on the basis of:
[0109] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0110] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0111] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0112] The metric can be obtained in the form of the following formula:
[0113]
[0114] Wherein, P(n) is an element associated with the first input signal [e.g., P1(n) or an element in the first set of normalized amplitude values], and Q(n) is an element associated with the second input signal [e.g., P2(n) or an element in the second set of normalized amplitude values].
[0115] The metric can be obtained in the form of the following formula:
[0116]
[0117] Wherein, P1(n) is an element associated with the first input signal [e.g., an element in the first set of normalized amplitude values of P1(n)], and P2(n) is an element associated with the second input signal [e.g., an element in the second set of normalized amplitude values].
[0118] The device can use a sliding window to perform optimization [for example, optimization can take into account the TD samples in the last 0.1s...1.0s].
[0119] The device can transform the information associated with the obtained first and second output signals (S'0, S'1) into the frequency domain.
[0120] The device can encode information associated with the obtained first and second output signals (S'0, S'1).
[0121] The device can store information associated with the obtained first and second output signals (S'0, S'1).
[0122] The device can send information associated with the obtained first and second output signals (S'0, S'1).
[0123] The apparatus of any of the foregoing technical solutions may include at least one of a first microphone (mic0) for obtaining a first input signal [M0] and a second microphone (mic1) for obtaining a second input signal [M1]. [For example, at a fixed distance]
[0124] An apparatus for teleconference may be provided, including the means as described above and equipment for transmitting information associated with the obtained first and second output signals (S'0, S'1).
[0125] A binaural system is disclosed, including the device described above.
[0126] An optimizer for iteratively optimizing physical parameters associated with a physical signal is disclosed, wherein the optimizer is configured to randomly generate a current candidate vector at each iteration to evaluate whether the current candidate vector performs better than the current best candidate vector.
[0127] The optimizer is configured to evaluate an objective function that is associated with the similarity or dissimilarity between the physical signals associated with the current candidate vector.
[0128] The optimizer is configured such that if the current candidate vector reduces the objective function relative to the current best candidate vector, the current candidate vector becomes the new current best candidate vector.
[0129] Physical signals can include audio signals obtained from different microphones.
[0130] Parameters may include delay and / or scaling factors for the audio signal obtained at a particular microphone.
[0131] The objective function is the Kullback-Leibler divergence. The Kullback-Leibler divergence can be applied to a first set and a second set of normalized magnitude values.
[0132] The objective function can be obtained by summing multiple products between the following two [for example, as in formula (6) or (8)]:
[0133] -The corresponding element in the first set of normalized amplitude values [P]i (n), where i is 0 or 1 [for example, as in formula (7)].
[0134] [China], and
[0135] - The logarithm of the quotient formed based on the following:
[0136] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0137] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0138] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0139] The objective function can be obtained as follows:
[0140] or
[0141]
[0142] Wherein, P1(n) or P(n) is an element associated with the first input signal [e.g., an element in the first set of normalized amplitude values, P1(n)], and P2(n) or Q(n) is an element associated with the second input signal.
[0143] According to one example, a method is provided for obtaining multiple output signals [S'0, S'1] associated with said different sound sources [source0, source1] based on multiple input signals [e.g., microphone signals] [M0, M1] which combine signals from different sound sources [source0, source1].
[0144] The method includes:
[0145] The first input signal [M0] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a1·z] of the second input signal [M1]. -d1 The combination of [·M1] [e.g., by subtracting a delayed and scaled version of the second input signal from the first input signal, e.g., by S'0 = M0(z) - a1·z] -d1 ·M1(z)], to obtain the first output signal [S'0];
[0146] The second input signal [M1] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a0·z] of the first input signal [M0]. -d0 The combination of [·M0] [e.g., by subtracting a delayed and scaled version of the first input signal from the second input signal, e.g., by S'1 = M1(z) - a0·z] -d0 ·M0(z)], to obtain the second output signal [S'1];
[0147] Determined using stochastic directional optimization [e.g., by performing one of the operations defined in other implementations; and / or by finding delay and decay values that minimize the objective function, which may be, for example, the objective function in equations (6) and / or (8)]:
[0148] A first scaling value [a0] is used to obtain a delayed and scaled version [a0*z] of the first input signal [M0]. -d0 *M0];
[0149] A first delay value [d0] is used to obtain a delayed and scaled version [a0*z] of the first input signal [M0]. -d0 *M0];
[0150] The second scaling value [a1] is used to obtain a delayed and scaled version [a1*z] of the second input signal [M1]. -d1 *M1];
[0151] as well as
[0152] The second delay value [d1] is used to obtain the delayed and scaled version [a1*z] of the second input signal. -d1 *M1].
[0153] According to one example, a method is provided for obtaining multiple output signals [S'0, S'1] associated with said different sound sources [source1, source2] based on multiple input signals [e.g., microphone signals] [M0, M1] which combine signals from different sound sources [source1, source2].
[0154] The method includes:
[0155] The first input signal [M0] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a1*z] of the second input signal [M1]. -d1The *M1] combination is used to obtain a first output signal [S'0], wherein the method is configured to apply a fractional delay [d1] to the second input signal [M1], wherein the fractional delay (d1) can indicate the signal (H) arriving from the second source (source1) to the first microphone (mic0). 1,0 The delay of *S1) (e.g., caused by H) 1,0 The delay (represented by the signal) and the signal (H) from the second source (source1) to the second microphone (mic1) 1,1 The delay of *S1) (e.g., caused by H) 1,1 The relationship and / or difference between the delays (represented by the fractional delay d1) and the approximate fraction H. 1,0 (z) / H 1,1 [The exponent of the z-term of the result of (z)];
[0156] The second input signal [M1] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a0*z] of the first input signal [M0]. -d0 The first input signal [M0] is combined to obtain a second output signal [S'1], wherein the method is configured to apply a fractional delay [d0] to the first input signal [M0]. The fractional delay (d0) can indicate the signal (H) arriving at the first microphone (mic0) from the first source (source0). 0,0 The delay of *S0) (e.g., caused by H) 0,0 The delay (represented by the signal) and the signal (H) from the first source (source0) to the second microphone (mic1) 0,1 The delay of *S0) (e.g., caused by H) 0,1 The relationship and / or difference between the delays (represented by the fractional delay) [In the example, the fractional delay d0 can be understood as the approximate fraction H]. 0,1 (z) / H 0,0 [The exponent of the z-term of the result of (z)];
[0157] Optimization was used to determine:
[0158] A first scaling value [a0] is used to obtain a delayed and scaled version [a0*z] of the first input signal [M0]. -d0 *M0];
[0159] A first fractional delay value [d0] is used to obtain a delayed and scaled version of the first input signal [M0].
[0160] [a0*z -d0 *M0];
[0161] The second scaling value [a1] is used to obtain a delayed and scaled version [a1*z] of the second input signal [M1].-d1 *M1];
[0162] as well as
[0163] The second fractional delay value [d1] is used to obtain a delayed and scaled version of the second input signal [M1].
[0164] [a1*z -d1 *M1].
[0165] According to one example, a method is provided for obtaining multiple output signals [S'0, S'1] associated with said different sound sources [source0, source1] based on multiple input signals [e.g., microphone signals] [M0, M1] that combine signals from different sound sources [source0, source1] [e.g., by subtracting a delayed and scaled version of the second input signal from a first input signal, and / or by subtracting a delayed and scaled version of the first input signal from a second input signal].
[0166] The first input signal [M0] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a1*z] of the second input signal [M1]. -d1 The first output signal [S'0] is obtained by combining [*M1] [for example, by subtracting a delayed and scaled version of the second input signal from the first input signal],
[0167] The second input signal [M1] or its processed (e.g., delayed and / or scaled) version is compared with the delayed and scaled version [a0*z] of the first input signal [M0]. -d0 The second output signal [S'1] is obtained by combining [*M0] [for example, by subtracting a delayed and scaled version of the first input signal from the second input signal],
[0168] Summing multiple products between the following two [for example, as in formula (6) or (8)]:
[0169] -The corresponding element in the first set of normalized amplitude values [P] i (n), where i is 0 or 1 [for example, as in formula (7)].
[0170] [China], and
[0171] - The logarithm of the quotient formed based on the following:
[0172] The corresponding element [P(n) or P1(n)] in the first set of normalized amplitude values; and
[0173] The corresponding element [Q(n) or Q1(n)] in the second set of normalized amplitude values,
[0174] In order to obtain the similarity [or dissimilarity] value between the signal portion [s0'(n)] described by the first set [P0(n) for n = 1 to ...] and the signal portion [s1'(n)] described by the second set [P1(n) for n = 1 to ...] [DKL(P||Q) or D(P0,P1) in formula (6) or (8).
[0175] According to one example, a method is provided in any of the foregoing method implementations, the method being configured to use the equipment as described above or as follows.
[0176] A non-transitory storage unit for storing instructions, which, when executed by a processor, cause the processor to perform a method according to any of the foregoing method implementation schemes. Attached Figure Description
[0177] Figure 1 The layout of the microphone and source is shown to aid in understanding the present invention;
[0178] Figure 2a The operating technology according to the present invention is illustrated;
[0179] Figure 2b The signal block diagram of convolution mixing and the mixing process is shown;
[0180] Figure 3 The performance evaluation of the BSS algorithm applied to simulated data is shown;
[0181] Figure 4 The layout of the microphone and sound source is shown to aid in understanding the present invention;
[0182] Figure 5 An apparatus according to the invention is shown;
[0183] Figure 6a , Figure 6b and Figure 6c This illustrates the results that can be obtained using the present invention; and
[0184] Figure 7 It shows Figure 5 The components of the device. Detailed Implementation
[0185] It is understood that by applying techniques such as those discussed above and below, signals can be processed to obtain multiple signals S1′ and S0′ that are separate from each other. As a result, the output signal S1′ is unaffected by the sound S0 (or has negligible or minimal effect), while the output signal S0′ is unaffected by the effect of the sound S1 on the microphone mic0 (or has minimal or negligible effect).
[0186] Example by Figure 2b Provided is a physical model showing the relationship between the generated sounds S0 and S1 and the signal 502 obtained by combining them from microphones M0 and M1. These results are expressed here using the z-transform (not indicated in some cases for brevity). As can be seen from block 501, the sound signal S0 undergoes a transfer function H 0,0 (z), and with (via transfer function H) 1,0 The modified (z) sound signal S1 is added together. Accordingly, signal M0(z) is obtained at microphone mic0, and the sound signal S1(z) is unintentionally included in the calculation. Similarly, signal M1(z) obtained at microphone mic1 includes the component associated with the sound signal S1(z) (via the transfer function H). 1,1 (z) is obtained) and obtained from the sound signal S0(z) (after passing through the transfer function H 0,1 The second undesirable component caused by (z) is the two. This phenomenon is called crosstalk.
[0187] To compensate for crosstalk, the solution indicated at block 510 can be used. Here, the multi-channel output signal 504 includes:
[0188] The first output signal S0′(z) (representing the sound S0 collected at microphone mic0 but with crosstalk removed) includes at least two components:
[0189] Input signal M0, and
[0190] The subtraction component 5031 (which is a delayed and / or scaled version of signal M1, and can be obtained by passing signal M1 through a transfer function) (to obtain)
[0191] • Output signal S1′(z) (representing the sound S1 collected at microphone mic1 after crosstalk has been removed), which includes:
[0192] Input signal M1, and
[0193] The subtraction component 5030 (which is a delayed and / or scaled version of the first input signal M0 obtained at microphone mic0, and which can be obtained by passing the signal M0 through a transfer function) (To obtain).
[0194] The mathematical explanation is provided below, but it is understood that at block 510, the subtraction components 5031 and 5030 compensate for the undesirable components caused at block 501. Thus, it is clear that block 510 allows for the (undesired) combination (510) of multiple (502) input signals [e.g., microphone signals] [(M0, M1)] from sound sources (source0, source1) to obtain multiple (504) output signals (S'0, S'1) associated with different sound sources (source0, source1). Block 510 can be configured to combine a first input signal (M0) or its processed [e.g., delayed and / or scaled] version with a delayed and scaled version (5031) of a second input signal (M1). -d1 Combining M1] (510) [for example by subtracting the delayed and scaled version of the second input signal from the first input signal, for example by S'0(z)=M0(z)-a1·z] -d1 ·M1(z)] to obtain a first output signal (S'0); wherein, the block is configured to combine the second input signal (M1) or its processed [e.g., delayed and / or scaled] version with the delayed and scaled version (5030)[a0·z] of the first input signal [M0]. -d0 Combining M0] (510) [for example, by subtracting a delayed and scaled version of the first input signal from the second input signal, for example by S'1(z) = M1(z) - a0·z] -d0 ·M0(z)] to obtain the second output signal [S'1].
[0195] While the z-transform is particularly useful in this case, other types of transforms can also be used or it can be operated on directly in the time domain.
[0196] Essentially, it can be understood that a pair of scaling values a0 and a1 corrects the amplitudes of the subtraction components 5031 and 5030 to obtain a scaled version of the input signal, and delays d0 and d1 can be understood as fractional delays. In the example, the fractional delay d0 can be understood as an approximate fraction H. 0,1 (z) / H 0,0 The exponent of the z-term of the result [z] is the fractional delay d1, which can indicate the signal (H) from the second source (source1) to the first microphone (mic0). 1,0 The delay of ·S1) (e.g., caused by H) 1,0 The delay (represented by the signal) and the signal (H) from the second source (source1) to the second microphone (mic1) 1,1 The delay of ·S1) (e.g., caused by H)1,1 The relationship and / or difference between the delays (represented by the fractional delay d1). In the example, the fractional delay d1 can be understood as the approximate fraction H. 1,0 (z) / H 1,1 The exponent of the z-term of the result [z] is the fractional delay d0, which can indicate the signal (H) from the first source (source0) to the first microphone (mic0). 0,0 The delay of ·S0) (e.g., caused by H) 0,0 The delay (represented by the signal) and the signal (H) from the first source (source0) to the second microphone (mic1) 0,1 The delay of ·S0) (e.g., caused by H) 0,1 The relationship and / or difference between the delays (represented by the fractional delay d0) [in the example, the fractional delay d0 can be understood as the approximate fraction H] 0,1 (z) / H 0,0 The exponent of the z term of the result (z)].
[0197] As will be explained later, the optimal value can be found (also indicated by reference number 564), in particular:
[0198] • A first scaling value [a0], for example, which is used to obtain a delayed and scaled version 5030 [a0·z] of the first input signal [502, M0]. -d0 ·M0];
[0199] • A first fractional delay value [d0], for example, which is used to obtain a delayed and scaled version 5030 [a0·z] of the first input signal [502, M0]. -d0 ·M0];
[0200] • A second scaling value [a1], for example, which is used to obtain a delayed and scaled version 5031 [a1·z] of the second input signal [502, M1]. -d1 ·M1]; and
[0201] • The second fractional delay value [d1], for example, is the delayed and scaled version 5031 [a1·z] used to obtain the second input signal [502, M1]. -d1 ·M1).
[0202] The techniques used to obtain the optimal scaling values a0 and a1, as well as the delay values d0 and d1, are discussed here, with particular reference to Figure 5 For example from Figure 5 As can be seen, a stereo or multi-channel signal 502 (including input signals M0(z) and M1(z)) is obtained. It can also be seen that this method can be iterative, in this sense, iteratively looping multiple times to obtain the optimal values for the scaling and delay to be used.
[0203] Figure 5 Output 504 is shown, which is formed by signals S′0(z) and S′1(z), and these signals are optimized, for example, after multiple iterations. Figure 5 The mixing block 510 is shown, which can be Figure 2b Block 510.
[0204] The multi-channel signal 512 (including its channel components, i.e., multiple input signals S′0(z) and S′1(z)) is thus obtained by utilizing scaling values a0 and a1 and delay values d0 and d1, which are optimized more and more along the iteration.
[0205] In block 520, normalization is performed on signals S′0(z) and S′1(z). An example of normalization is provided by equation (7), expressed as the following quotient:
[0206]
[0207] Here, i = 0, 1, indicating the existence of a normalized value P0(n) for input signal M0 and a normalized value P1(n) for input signal M1. Index n is the time index of the time-domain input signal. Here, s′ i (n) is the signal M i The time-domain sample index of (i = 0, 1) (it is not the z-transform). |s′ i (n)| indicates the obtained s′ i The magnitude (e.g., absolute value) of (n) and therefore positive, or at worst 0. This implies that the numerator in formula (7) is positive, or at worst 0. ‖|s′i|‖1 indicates that the denominator in formula (7) is derived from the vector s′ i The 1-norm is formed. The 1-norm ||…|||1 indicates the magnitude |s′ i The sum of (n)|, where n traverses the signal samples, e.g., up to the current index (e.g., signal samples can be extracted within a predetermined window from past indices to the current index). Therefore, ||s′| i |‖1 (which is the denominator in formula (7)) is positive (or 0 in some cases). Furthermore, |s′i(n)|≤‖|s′ i |‖1 always holds true, which implies that 0≤P i (n)≤1 (i=0,1). Furthermore, the following formula was also verified:
[0208]
[0209] It has thus been noted that P0(n) and P1(n) can be artificially considered as probabilities, because they are verified by using equation (7):
[0210] 1.
[0211] 2.
[0212] Where i = 0, 1 (further discussion follows). “∞” is used in mathematical form, but can be approximated over the signal under consideration.
[0213] Note that other types of normalization can be provided, not just those obtained through formula (7).
[0214] Figure 5 Block 530 is shown, which takes a normalized value 522 as input and outputs a similarity value (or dissimilarity value) 532 to provide information about the relationship between the first input value M0 and the second input value M1. Block 530 can be understood as a block that measures a metric that indicates the degree of similarity (or dissimilarity) between the input signals M0 and M1.
[0215] It is understood that the metric chosen to indicate the similarity or dissimilarity between the first and second input values can be the so-called Kullback-Leibler divergence (KLD). This can be obtained using formula (6) or (8):
[0216]
[0217] A discussion is now provided on how to obtain the Kullback-Leibler divergence (KLD). Figure 7 It shows Figure 5 An example of a downstream block 530 of block 520. Block 520 thus provides P0(n) and P1(n) (522), for example using formula (7) as discussed above (other techniques may be used). Block 530 (which can be understood as a Kullback-Leibler processor or a KL processor) can be adapted to obtain metric 532, in this case, metric 532 is the Kullback-Leibler divergence as calculated in formula (8).
[0218] refer to Figure 7 In the first branch 700a, the quotient 702' between P0(n) and P1(n) is calculated in block 702. In block 706, the logarithm of the quotient 702' is calculated to obtain the value 706'. The logarithmic value 706' can then be used to scale the normalized value P0 in scaling block 710 to obtain the product 710'. In the second branch 700b, the quotient 704' is calculated in block 704. The logarithm 708' of the quotient 704' is calculated in block 708. The logarithmic value 708' is then used to scale the normalized value in scaling block 712 to obtain the product 712'.
[0219] In adder block 714, values 710' and 712' (obtained respectively in branches 700a and 700b) are combined with each other. The combined value 714' is summed with each other along the sample domain indices in block 716. The summed value 716' can be reversed in block 718 (e.g., by scaling by -1) to obtain the reversed value 718'. It should be noted that value 716' can be interpreted as a similarity value, while the reversed value 718' can be interpreted as a dissimilarity value. As explained above, either value 716' or value 718' can be provided to optimizer 560 as a metric 532 (value 716' indicates similarity, value 718' indicates dissimilarity).
[0220] Therefore, optimizer block 530 can thus allow equation (8) to be achieved, i.e. To achieve formula (6), for example D KL It can be easily obtained from Figure 7 Remove blocks 704, 708, 712, and 714 from the list, and replace P0 with P and P1 with Q.
[0221] Kullback-Leibler divergence was originally conceived to provide measurements regarding probabilities and, in principle, is independent of the physical meaning of the input signals M0 and M1. Nevertheless, we have already understood that by normalizing the signals S′0 and S′1 and obtaining normalized values such as P0(n) and P1(n), Kullback-Leibler divergence provides an effective measure for measuring the similarity / dissimilarity between input signals M0 and M1. Thus, the normalized amplitude values of the time-domain samples can be regarded as a probability distribution, and this measure can then be measured (e.g., as Kullback-Leibler divergence, for example, obtained by formula (6) or (8)).
[0222] Now refer to again Figure 5 For each iteration, metric 532 provides a good estimate of the effectiveness of the scaling values a0 and a1, and the delay values d0 and d1. Along the iterations, different candidate values for scaling values a0 and a1, and delay values d0 and d1, will be selected from those candidate values that present the lowest similarity or the highest dissimilarity.
[0223] Block 560 (the optimizer) takes metric 532 as input and outputs 564 (vectors) of candidates for delay values d0 and d1 and scaling values a0 and a1. The optimizer 560 can measure different metrics obtained for different candidate groups a0, a1, d0, d1, modify them, and select the candidate group associated with the lowest similarity (or highest dissimilarity) 532. Therefore, the output 504 (output signals S′0(z), S′1(z)) will provide the best approximation. The candidate values 564 can be grouped into vectors, which can then be refined, for example, through randomization techniques (…). Figure 5 A random generator 540 is shown that provides random input 542 to optimizer 560. Optimizer 560 can utilize weights to scale candidate values 564 (a0, a1, d0, d1) by weights (e.g., randomly). Initial coefficient weights 562 can be provided, for example, by default. An example of the processing of optimizer 560 is provided and discussed extensively below (“Algorithm 1”). Figure 5 The diagram also shows the lines of the algorithm and... Figure 5 Possible correspondences between components.
[0224] As can be seen, optimizer 564 outputs a vector 564 of values a0, a1, d0, d1, which is subsequently reused in mixing block 510 to obtain new values 512, new normalized values 522, and new metric 532. After a certain number of iterations (which may be, for example, predefined), the maximum number of iterations may be, for example, a number chosen between 10 and 20. Essentially, optimizer 560 can be understood as finding the delay and iteration values that minimize the objective function, which may be, for example, the metric 532 obtained in block 530 and / or using formulas (6) and (8).
[0225] Therefore, it is understood that optimizer 560 can be based on random direction optimization techniques to form candidate vectors of candidate parameters [e.g., having four elements, such as 564, a0, a1, d0, d1], wherein the candidate vectors are iteratively refined by modifying the candidate vectors in random directions.
[0226] Essentially, the candidate vector (indicating the ordered values of a0, a1, d0, d1) can be iteratively refined by modifying the candidate vector in a random direction. For example, after a random input of 542, different candidate values can be modified by using different weights that vary randomly. The random direction can mean, for example, that some candidate values increase while others decrease, or vice versa; there is no predefined rule. Similarly, the increment of the weights can be random, even if the maximum threshold is predefined.
[0227] The optimizer 460 can form a candidate vector of candidate parameters [a0, a1, d0, d1], wherein, for each iteration, the candidate vector is perturbed [e.g., randomly] by applying a vector of uniformly distributed random numbers [e.g., each between -0.5 and +0.5], which is element-wise multiplied (or added) with the elements of the candidate vector. Gradient processing, for example, can be avoided by using randomized direction optimization. Thus, by randomly perturbing the vector of coefficients, optimal values for a0, a1, d0, d1 can be obtained progressively, resulting in an output signal 504 in which the combined sounds S0 and S1 have been appropriately compensated. This algorithm will be discussed in detail below.
[0228] In this example, the multichannel input signal 502, which is formed by two input channels (e.g., M0, M1), is always referenced. However, the same embodiment described above also applies to more than two channels.
[0229] In these examples, the logarithm can use any base. One can imagine that the base discussed above is 10.
[0230] Detailed discussion of the technology
[0231] One objective is a system for teleconferencing to separate two speakers or a speaker from an instrument or noise source, in a small office environment, not too far from a stereo microphone in a stereo webcam, for example. The speaker or source is assumed to be on opposite (left and right) sides of the stereo microphone. For usability in real-time teleconferencing, we want it to work online with the lowest possible latency. For comparison, this paper focuses on an offline implementation. The proposed method works in the time domain, using an attenuation factor and fractional delay between microphone signals to minimize crosstalk, based on the principle of a fractional delay summation beamformer. This method has the advantage of requiring fewer variables to optimize compared to other methods, and avoids the permutation problem of ICA-like methods in the frequency domain. To optimize the separation, we minimize a negative Kullback-Leibler derived objective function between the resulting separated signals. For optimization, we use a novel “random direction” algorithm that does not require gradients, and is very fast and robust. We evaluate our method on a convolutional mixture of speech signals obtained from the TIMIT dataset using a room impulse response simulator and real recordings. The results show that our method is competitive in terms of separation performance for the proposed scenario, and has lower computational complexity and system latency compared with existing methods.
[0232] Index words—Blind source separation, time domain, binaural room impulse response, optimization
[0233] 1. Introduction, Previous Methods
[0234] Our system is suitable for situations with two microphones where it is desirable to separate the two audio sources. This could be, for example, a telephone conference scenario in an office with a stereo webcam and two speakers around it, or for hearing aids, where low computational complexity is important.
[0235] Previous methods: An earlier method was Independent Component Analysis (ICA). It could unmix signals without delay in the mixture. It found the coefficients of the unmixing matrix by maximizing non-gaussianity or maximizing the Kullback-Leibler divergence [1, 2]. However, for audio signals and stereo microphone pairs, there is always a propagation delay, usually convolved with the room impulse response [3] in the mixture. The way to deal with this is often to apply a short-time Fourier transform (STFT) [4] to the signal, such as AuxIVA [5] and ILRMA [6, 7]. This converts the signal delay into a complex factor in the STFT subband, and (complex) ICA can be applied to the resulting subband (e.g. [8]).
[0236] Problem: The issue here is the arrangement of subbands. Separated sources may appear in different subbands in different orders; the gain for different sources in different subbands may be different, resulting in modified spectral shape and spectral flattening. Similarly, there is a signal delay when applying STFT. It requires assembling the signal into blocks, which requires a system delay corresponding to the block size [9, 10].
[0237] Temporal methods, such as TRINICON
[11] , or methods using STFTs with short blocks and more microphones[12, 13], have the advantage of not having the large blocking delay of STFTs, but generally have high computational complexity, making them difficult to use on small devices.
[0238] See also Figure 1 This shows the speaker and microphone setup in the simulation.
[0239] 2. Proposed methods
[0240] To avoid the processing delays associated with frequency-domain methods, time-domain methods are used. Instead of FIR filters, we use IIR filters, which are implemented as fractional-delay all-pass filters [14, 15], based on the principles of attenuation factors, fractional-delay summation, or adaptive beamforming [16, 17, 18]. This has the advantage that each of these filters has only two coefficients: fractional delay and attenuation. For the 2-channel stereo case, this results in only four coefficients in total, which are easier to optimize. For simplicity, we also omit dreaving and focus on minimizing crosstalk. In effect, we model the relative transfer function between the two microphones using attenuation and pure fractional delay. We then apply a novel “random direction” optimization, similar to “differential evolution”.
[0241] We assume a mixture recorded from two sound sources (S0 and S1) using two microphones (M0 and M1). However, the same result applies to more than two sources. Figure 1 As shown, it can be assumed that the sound source is in a fixed position. To avoid the need for modeling non-causal impulse responses, the sound source must be on different half-planes (left and right) of the microphone pair.
[0242] Instead of the commonly used STFT, we can use the z-transform for mathematical derivation because it does not require decomposing the signal into blocks with associated delays. This makes it suitable for time-domain implementations without algorithmic delays. Remember, for sample index n, the (one-sided) z-transform of the time-domain signal x(n) is defined as... We use uppercase letters to represent z-transform domain signals.
[0243] We define s0(n) and s1(n) as two time-domain sound signals at time (sample index) n, and define their z-transforms as S0(z) and S1(z). The two microphone signals (together indicated by 502) are m0(n) and m1(n), and their z-transforms are M0(z) and M1(z) (Figure 2).
[0244] The room impulse response (RIR) from source i to microphone j is h. i,j (n), and their z-transformation is H i,j (z). Therefore, our convolutional hybrid system can be described in the z-domain as:
[0245]
[0246] In simplified matrix multiplication, we rewrite equation (1) as:
[0247] M(z)=H(z)·S(z) (2)
[0248] For ideal sound source separation, we need to invert the mixing matrix H(z). Therefore, our sound source can be calculated as:
[0249]
[0250] Since det(H(z)) and the diagonal elements of the inverse matrix are linear filters that do not contribute to demixing, we can ignore them during separation and substitute them into the left side of equation (3). This yields
[0251]
[0252] in, and This is now a relative room transfer function.
[0253] Next, we use a delay of d i Fractional delay and decay factor a for each sample i To approximate these relative room transfer functions,
[0254]
[0255] Where i,j∈{0,1}.
[0256] This approximation works particularly well when the reverberation or echo is not too strong. For a delay of d... i The fractional delay of each sample is used in the next section (2.1), where we will employ a fractional delay filter. Note that, for simplicity, the fractional delay will be calculated from the determinant and the matrix diagonal H. i,i The linear filter obtained by (z) is retained on the left side, which means that there is no de-reverberation.
[0257] Reference here Figure 1 An example is provided: Assume two sources in a free field, with no reflections, symmetrically positioned on opposite sides of a stereo microphone pair. The distance (in samples) of source0 to the nearest microphone should be m0 = 50 samples, and the distance to the far microphone should be m1 = 55 samples. The sound amplitude will be determined by a function with a certain constant k. Attenuation. Then the room transfer function in the z-domain is H. 0,0 (z)=H 1,1 (z)=k / 50 2 ·z -50 and And a relative room transfer function is The other relative room transfer function is the same. We see that, in this simple case, the relative room transfer function is indeed 0.825·z. -5 More precisely, it refers to attenuation and delay. Figure 2bThe signal flow diagrams for the convolution mixing and demixing processes can be seen in the diagram. Figure 2b (Signal block diagram of the convolution mixing and demixing process).
[0258] 2.1. Fractional Delay All-Pass Filter
[0259] Delay used to implement equation (5) The fractional-delay all-pass filter plays a crucial role in our scheme because it generates the IIR filter with only a single coefficient and allows for the implementation of an exact fractional delay (where the exact delay is not an integer value), which is necessary for good crosstalk cancellation. The IIR properties also allow for efficient implementation methods. [14, 15] describe a practical approach for designing fractional-delay all-pass filters based on IIR filters with a maximum flat group delay
[19] . For a fractional delay τ = d i We use the following equation to obtain the coefficients for the fractional-delay all-pass filter. Its transfer function in the z-domain is A(z), see
[14] .
[0260]
[0261] Where D(z) is Order, defined as:
[0262]
[0263] The filter d(n) is generated as follows:
[0264] d(0)=1,
[0265]
[0266] For 0 ≤ n ≤ (L-1).
[0267] 2.2. Objective Function
[0268] As the objective function, we use the function D(P0,P1) derived from the Kullback-Leibler divergence (KLD).
[0269]
[0270] Where P(n) and Q(n) are the probability distributions of our (unmixed) microphone channels, and n is a discrete distribution that traverses the spectrum.
[0271] To make the calculation faster, we avoid calculating the histogram. Instead, we use the normalized amplitude of the time-domain signal itself.
[0272]
[0273] Here, n is now the time-domain sample index. Note that P i (n) has properties similar to probability, namely:
[0274] 1.
[0275] 2.
[0276] Where i = 0, 1. Instead of directly using the Kullback-Leibler divergence, we use the summation of D... KL (P‖Q)+D KL (Q‖P) transforms the objective function into a symmetric (distance) function, as this makes the separation between the two channels more stable. To apply minimization instead of maximization, we take its negative value. Therefore, the resulting objective function D(P0,P1) is:
[0277]
[0278] 2.3. Optimization
[0279] A widely used optimization method for BSS is gradient descent. This method has the advantage of finding the “steepest” way to reach the optimum, but it requires calculating gradients and is easily trapped in local minima or slowed down by “valleys” in the objective function. Therefore, to optimize the coefficients, we use a novel “random direction” optimization, similar to “differential evolution” [20, 21, 22]. Instead of using the difference of the coefficient vectors for updates, we use a weight vector to model the expected variance distribution of the coefficients. This results in a very simple and very fast optimization algorithm that can also be easily applied to real-time processing, which is important for real-time communication applications. The algorithm starts with a fixed starting point [1.0, 1.0, 1.0, 1.0], which we find leads to robust convergence behavior. The current point is then perturbed by element-wise multiplying the weight vector (line 10 in Algorithm 1) by a vector of uniformly distributed random numbers between -0.5 and +0.5 (random direction). If the perturbed point has a lower objective function value, it is chosen as the next current point, and so on. The pseudocode for this optimization algorithm can be seen in Algorithm 1. Here, minabskl_i (indicated as negabskl_i in Algorithm 1) is our objective function, which calculates the KLD based on the coefficient vector coeffs and the microphone signals in array X.
[0280] We found that 20 iterations (and therefore only 20 objective function evaluations) are sufficient for the test file (the entire file each time), which makes the algorithm very fast.
[0281] This optimization can be performed, for example, in block 560 (see above).
[0282] Algorithm 1 is shown below.
[0283]
[0284] 3. Experimental Results
[0285] In this section, we evaluate the proposed time-domain separation method, which we call AIRES (time domain fractional delay separation), by using simulated room impulse responses and different speech signals from the TIMIT dataset
[23] . In addition, real-life recordings were made in a real room environment. To evaluate the performance of our proposed method, it was compared with state-of-the-art BSS algorithms, namely time domain TRINICON
[11] , frequency domain AuxIVA [5], and ILRMA [6, 7]. The implementation of TRINICON BSS has been obtained from its authors. The implementations of AuxIVA and ILRMA BSS are taken from
[24] and
[25] , respectively. The experiments were performed using MATLAB R2017a on a laptop computer with an 8th generation i7 CPU core and 16Gb RAM.
[0286] 3.1. Utilizing the separation properties of synthetic RIR
[0287] A room impulse response simulator based on image modeling techniques [26, 27] was used to generate the room impulse response. For the simulation setup, the room size was chosen to be 7m × 5m × 3m. The microphone was placed in the center of the room at [3.475, 2.0, 1.5]m and [3.525, 2.0, 1.5]m, with a sampling frequency of 16kHz. Ten pairs of speech signals were randomly selected from the entire TIMIT dataset and convolved using a simulated RIR. For each pair of signals, at a random angular position of the sound source relative to the microphone, at 4 different distances and 3 reverberation times (RT). 60 The simulation was repeated 16 times. Table 1 shows the common parameters used in all simulations, and the settings can be visualized. Figure 1 The evaluation of separation performance was objectively accomplished by calculating the signal distortion rate (SDR)
[28] when the original speech source was available, as well as the computation time. The results were presented in... Figure 3 As shown in the image.
[0288] The results show that our method performs well for reverberation times less than 0.2 s. For RT 60 =0.05s, the average SDR measurement across all distances is 15.64dB, and for RT60 =0.1s, which is 10.24dB. For the reverberation time RT 60 =0.2s, the proposed BSS algorithm ranks second after TRINICON and ILRMA. The average computation time on all simulations (on our computer) is shown in Table 2. As can be seen, AIRES outperforms all current state-of-the-art algorithms in terms of computation time.
[0289] Through listening to the results, we found that an SDR of about 8 dB achieved good speech intelligibility, and our method indeed did not produce unnatural echo artifacts.
[0290] Table 1: Parameters used in the simulation:
[0291]
[0292] Figure 3 Characteristics (showing the performance evaluation of the BSS algorithm applied to simulated data):
[0293] (a)RT 60 =0.05s
[0294] (b)RT 60 =0.1s
[0295] (c)RT 60 =0.2s
[0296] Table 2: Comparison of average computation time (simulated dataset).
[0297] BSS AIRES TRINICON ILRMA AuxIVA Calculation time 0.05s 23.4s 4.4s 0.6s
[0298] 3.2. Real-life experiments
[0299] Finally, real-life experiments were conducted to evaluate the proposed sound source separation method. Real recordings were captured in three different room types: a small apartment room (3m×3m), an office (7m×5m), and a large conference room (15m×4m). For each room type, ten stereo mixes of two speakers were recorded. Since there is no “ground truth” signal, the mutual information metric between the separated channels was calculated to evaluate the separation performance
[29] . The results can be seen in Table 3. Note that the average mutual information of the mixed microphone channels is 1.37, and the lower the mutual information between the separated signals, the better the separation.
[0300] Table 3: Comparison of separation performance using average mutual information (MI, Mutual Information, real recording).
[0301] BSS AIRES TRINICO ILRMA AuxIVA Average MI 0.5 0.47 0.52 0.52 Average computation time 0.22s 120.2s 7.86s 1.5s
[0302] As can be seen from the comparison in Table 3, the performance trend of separating convolutional mixed tones from real recordings is consistent with that of simulated data. Therefore, it can be concluded that although AIRES is simple, it can compete with existing blind source separation algorithms.
[0303] 4. Conclusion
[0304] In this paper, a fast temporal blind source separation technique based on the estimation of the IIR fractional delay filter is presented to minimize crosstalk between two audio channels. It has been shown that the estimation of fractional delay and attenuation factor yields fast and efficient separation of source signals from stereo convolutional mix. To this end, we introduce an objective function derived from negative Kullback-Leibler divergence. To make the minimization robust and fast, we present a novel “stochastic direction” optimization method, which is similar to “differential evolution” optimization. A series of experiments were conducted to evaluate the proposed BSS technique. We evaluate and compare our system with other state-of-the-art methods based on simulated data and real room recordings. The results show that our system is competitive in its separation performance despite its simplicity, with much lower computational complexity and no system latency. Online adaptation for real-time minimum latency applications and mobile sources such as mobile speakers is also achieved. These properties make AIRES well-suited for real-time applications on small devices, such as hearing aids or small teleconference settings. Test procedures for AIRES BSS are available on our GitHub
[30] .
[0305] Further aspects (see also the examples above and / or below)
[0306] Further aspects are discussed here, such as methods for separating multi-channel or stereo audio sources and their update methods. It minimizes the objective function (e.g., mutual information) and uses crosstalk reduction by acquiring signals from other channels, applying attenuation factors and delays (which can be fractional delays), and subtracting them from the current channel, for example. For instance, it uses a "random direction" approach to update the delay and attenuation coefficients.
[0307] See also the following:
[0308] 1: A method for separating multi-channel or stereo sound sources, wherein the coefficients used for separation are iteratively updated by adding a random update vector, and if this update results in improved separation, it is retained in the next iteration; otherwise, it is discarded.
[0309] 2: According to the method in aspect 1, for each coefficient's random update, there is a suitable fixed variance.
[0310] 3: The method according to aspect 1 or 2, wherein the update occurs on a real-time audio stream, wherein the update is based on past audio samples.
[0311] 4: According to the method of aspect 1, 2 or 3, wherein the updated variance is gradually reduced over time.
[0312] Other aspects could be:
[0313] 1) A method for online optimization of parameters, which utilizes a multidimensional input vector x to minimize the objective function by first determining a test vector from a given neighborhood space, then updating the previous best vector with a search vector, and if the objective function is obtained in this way, the updated vector becomes the new best; otherwise, it is discarded.
[0314] 2) The method of aspect 1 is applied to the task of separating N sources in an audio stream with N channels, wherein the coefficient vector consists of the delay and attenuation values required to cancel unwanted sources.
[0315] 3) The method in aspect 2 is applied to tasks involving telephone conferences with more than one speaker at one or more sites.
[0316] 4) The method in aspect 2 is applied to the task of separating audio sources before encoding them with an audio encoder.
[0317] 5) The method in aspect 2 is applied to the task of separating speech sources from music or noise sources.
[0318] Additional aspects
[0319] Additional aspects are discussed here.
[0320] Introduction Stereo Separation
[0321] The goal is to separate the source using multiple microphones (two in this case). Different microphones pick up sounds with different amplitudes and delays. The following discussion considers programming examples in Python. This is to facilitate understanding, test whether and how the algorithm works, and assess the reproducibility of the results, making the algorithm testable and useful to other researchers.
[0322] Other aspects
[0323] Other aspects will be discussed below. There are not necessarily boundaries between them, but they can be combined to create new embodiments. Each point can be independent of the other points and can be combined alone or with other features (e.g., other points), or supplemented or further specified by other features discussed above or below, as well as some features disclosed in at least some of the examples and / or embodiments above and / or below.
[0324] Spatial perception
[0325] The ear primarily uses two effects to estimate the specific direction of sound:
[0326] Interaural level difference (ILD)
[0327] Interauricular Time Difference (ITD)
[0328] Music from the recording studio mostly uses only level differences (no time differences), a technique known as "panning".
[0329] - Signals from stereo microphones, such as those from stereo webcams, primarily show time differences and not much level difference.
[0330] Stereo recording using a stereo microphone
[0331] - Setting up a stereo microphone with a time difference.
[0332] -Observation: The effects of sound delay from finite sound speed and attenuation can be described by a mixing matrix with delay.
[0333] See Figure 2a Stereo conference call setup. Observe the signal delay between microphones.
[0334] Stereo separation, ILD
[0335] - Cases without time difference (translation only) are relatively easy to handle - typically, independent component analysis is used.
[0336] It calculates a "demixing" matrix that produces statistically independent signals.
[0337] - Remember: The joint entropy of two signals X and Y is H(X,Y).
[0338] -The conditional entropy is H(X│Y)
[0339] Statistical independence also means that: H(X│Y)=H(X), H(X,Y)=H(X)+H(Y)
[0340] Mutual information: I(X,Y)=H(X,Y)–H(X│Y)–H(Y│X)≥0,
[0341] Regarding independence, it becomes 0.
[0342] Typically, optimization minimizes this mutual information.
[0343] ITD's previous methods
[0344] ITD situations with time lags are more difficult to resolve.
[0345] The "minimum variance principle" attempts to suppress noise from the microphone array.
[0346] It estimates the correlation matrix of the microphone channels and applies principal component analysis (PCA) or Karhounen-Loeve transform (KLT).
[0347] This causes the resulting vocal tract to become decorrelated.
[0348] - The channel with the lowest intrinsic value is assumed to be noise and set to zero.
[0349] Previous methods of ITD
[0350] - For stereo, we only use two channels and calculate and play signals with larger and smaller intrinsic values: python KLT_separation.py
[0351] - Observation: The channel with the larger intrinsic value appears to have lower frequencies, while the other has (weaker) higher frequencies, which tend to be just noise.
[0352] ITD, a New Approach
[0353] - Approach: The demixing matrix requires the delay of the demixed signal to be used for inverting the mixing matrix.
[0354] - Important simplification: We assume that there is only a single signal delay path between the microphones used for each of the two sources.
[0355] In reality, there are multiple delays from room reflections, but these are usually much weaker than those from a direct path.
[0356] Stereo separation, ITD, a new approach
[0357] Assume the mixing matrix has a decay of *a* and a delay of *d* samples. In the z-transform domain, the delay of *d* samples is determined by z... -d The factor representation (observed that this delay can be a fractional sample):
[0358]
[0359] Therefore, the unmixing matrix is its inverse matrix:
[0360]
[0361] It turns out that we can simplify the unmixed matrix by discarding the fraction of the determinant at the front of the matrix without sacrificing actual performance.
[0362] Coefficient Calculation
[0363] The coefficients a and d need to be obtained through optimization. The coefficients can be found again by minimizing the mutual information of the resulting signals. For mutual information, joint entropy is required. In Python, it can be calculated from a two-dimensional probability density function using numpy.histogram2d. It can be called using histogram2d, xedges, yedges = np.histogram2d(x[:,0], x[:,1], bins = 100).
[0364] Optimization of the objective function
[0365] Note that our objective function may have several minima! This means that convex optimization cannot be used if its starting point is not sufficiently close to the global minimum.
[0366] Therefore, non-convex optimization is required, and Python has very powerful non-convex optimization methods:
[0367] scipy.optimize.differential_evolution . use
[0368] coeffs_minimized=opt.differential_evolution(mutualinfocoeffs,bounds,
[0369] Use args = (X,),tol = 1e-4,disp = True) to call it.
[0370] Other objective functions
[0371] However, this is computationally complex. Therefore, we need to find alternative objective functions with the same minimum value. We consider the Kullback-Leibler and Itakura-Saito divergences. The Kullback-Leibler divergence of two probability distributions P and Q is defined as:
[0372]
[0373] Where i is an ergodic (discrete) distribution. To avoid computing histograms, our trick is to simply treat the normalized magnitudes of the time-domain samples as a probability distribution. Since these are dissimilarity measures, they need to be maximized, and therefore their negative values need to be minimized.
[0374] Kullback-Leibler Python functions
[0375] In Python, this objective function is:
[0376] def minabsklcoeffs(coeffs,X):
[0377] #Calculate the normalized amplitude of the vocal tract, then apply
[0378] #Kullback-Leibler divergence
[0379] X_prime = unmixing(coeffs, X)
[0380] X_abs = np.abs(X_prime)
[0381] # Normalize it to sum() = 1, making it look like a probability:
[0382] X_abs[:,0]=X_abs[:,0] / np.sum(X_abs[:,0]
[0383] X_abs[:,1]=X_abs[:,1] / np.sum(X_abs[:,1])
[0384] # Print(“Kullback-Leibler Divergence calculation”)
[0385] abskl=np.sum(X_abs[:,0]*np.log((X_abs[:,0]+1e-6) / (X_abs[:,1]+1e-6))
[0386] →)
[0387] return –abskl
[0388] (Here, minabsklcoeffs corresponds to minabsk_i in Algorithm 1)
[0389] Comparison of objective functions
[0390] We fixed all the coefficients except one of the delay coefficients to compare the objective function in a single graph.
[0391] Figure 6a The objective function and example coefficients for the example signal are shown. It is observed that these functions do indeed have the same minimum value! “abskl” is the Kullback-Leibler term for the absolute value of the signal and is the smoothest.
[0392] Optimization Example
[0393] - Optimizations using mutual information: Python
[0394] ICAmutualinfo_puredelay.py
[0395] (This requires slower iterations of ca.121)
[0396] - Optimizations using Kullback-Leibler: python ICAabskl_puredelay.py
[0397] (This requires faster iterations of ca.a.39)
[0398] The signals obtained from listening confirm that they have indeed been separated clearly enough, whereas they were not before.
[0399] Further accelerate optimization
[0400] To further accelerate the optimization, convex optimization needs to be enabled. This requires an objective function with only one minimum (in the optimal case). Most of the smaller local minima arise from rapidly changing high-frequency components in our signal.
[0401] Method: The objective function is computed based on low-pass and downsampled versions of our signal (which further speeds up the computation of the objective function). The low-pass filter needs to be narrow enough to remove local minima, but also wide enough to obtain sufficiently accurate coefficients.
[0402] Further accelerate optimization, low-pass
[0403] We choose a downsampling factor of 8, and correspondingly choose a low-pass filter that covers 1 / 8 of the full bandwidth. We use the following low-pass filter, whose bandwidth is approximately 1 / 8 of the full bandwidth (1 / 8 of the Nyquist frequency).
[0404] Figure 6b The low-pass filter used and the amplitude-frequency response. The X-axis represents the normalized frequency, where the sampling frequency is 2π.
[0405] Objective function from low-pass filter
[0406] Now we can plot and compare our objective function again.
[0407] Figure 6c : Objective function and example coefficients for the low-pass filtered example signal. It is observed that these functions now indeed have primarily one minimum value!
[0408] "abskl" is again the smoothest (without small local minima).
[0409] Further accelerate optimization, for example
[0410] We can now try convex optimization, such as the conjugate gradient method.
[0411] -python ICAabskl_puredelay_lowpass.py
[0412] - It was observed that the optimization was completed almost instantly!
[0413] The resulting separation has the same quality as before.
[0414] Other explanations
[0415] The foregoing describes various examples and aspects of the inventive step. Similarly, further examples are defined by the appended claims (examples are also included in the claims). It should be noted that any example defined in the claims may be supplemented by any of the details (features and functionality) described in the preceding pages. Likewise, the examples described above may be used individually and may be supplemented by any of the features in another chapter or by any feature included in the claims.
[0416] The text within parentheses and square brackets is optional and defines further embodiments (those further defined by the claims). Similarly, it should be noted that aspects of the individuals described herein can be used individually or in combination. Therefore, details can be added to each aspect of the individual aspects, but not to the other aspects of the aspects.
[0417] It should also be noted that our content explicitly or implicitly describes the characteristics of the receivers of mobile communication devices and mobile communication systems.
[0418] Depending on the requirements of certain implementations, the examples can be implemented in hardware. Implementations can be executed using digital storage media with electronically readable control signals stored thereon, such as floppy disks, digital versatile optical discs (DVDs), Blu-ray discs, compact optical discs (CDs), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory, which cooperate (or are capable of cooperating with) a programmable computer system to enable the corresponding methods to be executed. Thus, the digital storage media can be computer-readable.
[0419] Typically, an example can be implemented as a computer program product having program instructions that, when run on a computer, are operable to perform one of the methods. The program instructions may, for example, be stored on a machine-readable medium.
[0420] Other examples include a computer program stored on a machine-readable medium for performing one of the methods described herein. In other words, examples of methods are thus computer programs having program instructions that, when run on a computer, are used to perform one of the methods described herein.
[0421] A further example of the method is thus a data carrier medium (or digital storage medium, or computer-readable medium) that includes a computer program recorded thereon for performing one of the methods described herein. The data carrier medium, digital storage medium, or recording medium is a tangible and / or non-transient signal, rather than an intangible and transient signal.
[0422] Further examples include processing units, such as computers or programmable logic devices, for performing one of the methods described herein.
[0423] Further examples include computers on which computer programs for performing one of the methods described herein are installed.
[0424] Further examples include apparatus or systems for transmitting (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, mobile device, memory device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.
[0425] In some examples, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some examples, the field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods can be implemented by any suitable hardware device.
[0426] The examples above are illustrative of the principles discussed herein. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent. Therefore, the intent is to be limited only by the scope of the following patent claims, and not by the specific details presented through the description and explanation of the examples herein.
[0427] References
[0428] [1]A. J.Karhunen,and E.Oja,“Independent component analysis.”John Wiley&Sons,2001.
[0429] [2]G.Evangelista,S.Marchand,M.D.Plumbley,and E.Vincent,“Sound sourceseparation,”in DAFX:Digital Audio Effects,second edition ed.John Wiley andSons,2011.
[0430] [3]J.Tariqullah,W.Wang,and D.Wang,“A multistage approach to blindseparation of convolutive speech mixtures,”in Speech Communication,2011,vol.53,pp.524–539.
[0431] [4]J.Benesty,J.Chen,and E.A.Habets,“Speech enhancement in the STFTdomain,”in Springer,2012.
[0432] [5]J. Z. and N.Ono,“A computationally cheaper method forblind speech separation based on auxiva and incomplete demixing transform,”inIEEE International Workshop on Acoustic Signal Enhancement(IWAENC),Xi’an,China,2016.
[0433] [6]D.Kitamura,N.Ono,H.Sawada,H.Kameoka,and H.Saruwatari,“Determinedblind source separation unifying independent vector analysis and nonnegativematrix factorization,”in IEEE / ACM Trans.ASLP,vol.24,no.9,2016,pp.1626–1641.
[0434] [7]“Determined blind source separation with independent low-rankmatrix analysis,”in Springer,2018,p.31.
[0435] [8]H.-C.Wu and J.C.Principe,“Simultaneous diagonalization in thefrequency domain(sdif)for source separation,”in Proc.ICA,1999,pp.245–250.
[0436] [9]H.Sawada,N.Ono,H.Kameoka,and D.Kitamura,“Blind audio sourceseparation on tensor representation,”in ICASSP,Apr.2018.
[0437]
[10] J.Harris,S.M.Naqvi,J.A.Chambers,and C.Jutten,“Realtimeindependent vector analysis with student’s t source prior for convolutivespeech mixtures,”in IEEE International Conference on Acoustics,Speech,andSignal Processing,ICASSP,Apr.2015.
[0438]
[11] H.Buchner,R.Aichner,andW.Kellermann,“Trinicon:A versatileframework for multichannel blind signal processing,”in IEEE InternationalConference on Acoustics,Speech,and Signal Processing,Montreal,Que.,Canada,2004.
[0439]
[12] J.Chua,G.Wang,andW.B.Kleijn,“Convolutive blind source separationwith low latency,”in Acoustic Signal Enhancement(IWAENC),IEEE InternationalWorkshop,2016,pp.1–5.
[0440]
[13] W.Kleijn and K.Chua,“Non-iterative impulse response shorteningmethod for system latency reduction,”in Acoustics,Speech and SignalProcessing(ICASSP),2017,pp.581–585.
[0441]
[14] I.Senesnick,“Low-pass filters realizable as all-pass sums:designvia a new flat delay filter,”in IEEE Transactions on Circuits and Systems II:Analog and Digital Signal Processing,vol.46,1999.
[0442]
[15] T.I.Laakso,V. M.Karjalainen,and U.K.Laine,“Splitting theunit delay,”in IEEE Signal Processing Magazine,Jan.1996.
[0443]
[16] M.Brandstein and D.Ward,“Microphone arrays,signal processingtechniques and applications,”in Springer,2001.
[0444]
[17] “Beamforming,”http: / / www.labbookpages.co.uk / audio / beamforming / delaySum.html,accessed:2019-04-21.
[0445]
[18] Benesty,Sondhi,and Huang,“Handbook of speech processing,”inSpringer,2008.
[0446]
[19] J.P.Thiran,“Recursive digital filters with maximally flat groupdelay,”in IEEE Trans.on Circuit Theory,vol.18,no.6,Nov.1971,pp.659–664.
[0447]
[20] S.Das and P.N.Suganthan,“Differential evolution:A survey of thestate-of-the-art,”in IEEE Trans.on Evolutionary Computation,Feb.2011,vol.15,no.1,pp.4–31.
[0448]
[21] R.Storn and K.Price,“Differential evolution-a simple andefficient heuristic for global optimization over continuous spaces,”inJournal of Global Optimization.11(4),1997,pp.341–359.
[0449]
[22] “Differential evolution,”http: / / www1.icsi.berkeley.edu / ~storn / code.html,accessed:2019-04-21.
[0450]
[23] J.Garofolo et al.,“Timit acoustic-phonetic continuous speechcorpus,”1993.
[0451]
[24] “Microphone array speech processing,”https: / / github.com / ZitengWang / MASP,accessed:2019-07-29.
[0452]
[25] “Ilrma,”https: / / github.com / d-kitamura / ILRMA,accessed:2019-07-29.
[0453]
[26] R.B.Stephens and A.E.Bate,Acoustics and Vibrational Physics,London,U.K.,1966.
[0454]
[27] J.B.Allen and D.A.Berkley,Image method for efficiently simulatingsmall room acoustics.J.Acoust.Soc.Amer.,1979,vol.65.
[0455]
[28] C.Fevotte,R.Gribonval,and E.Vincent,“Bss eval toolbox userguide,”in Tech.Rep.1706,IRISA Technical Report 1706,Rennes,France,2005.
[0456]
[29] E.G.Learned-Miller,Entropy and Mutual Information.Department ofComputer Science University of Massachusetts,Amherst Amherst,MA 01003,2013.
[0457]
[30] “Comparison of blind source separation techniques,”https: / / github.com / TUIlmenauAMS / Comparison-of-Blind-Source-Separation-techniques,accessed:2019-07-29.
Claims
1. Apparatus (500) for obtaining a plurality of output signals (504, S'0, S'1) associated with different sound sources (source0, source1) on the basis of a plurality of input signals (502, M0, M1) in which signals (S0, S1) from the different sound sources (source0, source1) are combined (501), wherein the apparatus being configured to combine (510) a first input signal (502, M0) or a processed version thereof with a delayed and scaled version (5031) of a second input signal (M1) to obtain a first output signal (504, S'0); wherein the apparatus is configured to combine (510) a second input signal (502, M1) or a processed version thereof with a delayed and scaled version (5030) of a first input signal (M0) to obtain a second output signal (504, S'1); wherein the apparatus is configured to determine, using a random direction optimization (560): a first scaling value (564, a0) for obtaining the delayed and scaled version (5030) of the first input signal (502, M0); a first delay value (564, d0) for obtaining the delayed and scaled version (5030) of the first input signal (502, M0); a second scaling value (564, a1) for obtaining the delayed and scaled version (5031) of the second input signal (502, M1); and a second delay value (564, d1) for obtaining the delayed and scaled version (5031) of the second input signal, wherein the random direction optimization (560) is such that a candidate parameter forms a candidate vector (564, a0, a1, d0, d1), wherein the candidate vector is iteratively refined by correcting the candidate vector in a random direction, wherein the random direction optimization (560) is such that a measure indicative of a similarity or dissimilarity between the first output signal and the second output signal is measured, and the first output signal and the second output signal are selected as those obtained using the candidate parameters associated with the measure indicative of the lowest similarity or the highest dissimilarity, wherein the measure (532) is processed as a Kullback-Leibler divergence.
2. The apparatus of claim 1, wherein, the delayed and scaled version (5031) of the second input signal (502, M1) to be combined with the first input signal (502, M0) is obtained by applying a fractional delay to the second input signal (502, M1).
3. The apparatus of claim 1, wherein, the delayed and scaled version (5030) of the first input signal (502, M0) to be combined with the second input signal (502, M1) is obtained by applying a fractional delay to the first input signal (502, M0).
4. The apparatus of claim 1, configured to sum (714, 716) a plurality of products (712', 710') between: - a respective element (P0) of a first set of normalized magnitude values; and - a logarithm (706') of a quotient (702') formed on the basis of: o the respective element (P0) of the first set of normalized magnitude values (522); and o a respective element (P1) of a second set of normalized magnitude values (522), to obtain (530) a value (D, 532) describing a similarity or dissimilarity between the signal portion (s'0) described by the first set of normalized amplitude values (P0) and the signal portion (s'1) described by the second set of normalized amplitude values (P1). KL , D, 532).
5. The apparatus of claim 1, wherein, The random direction optimization (560) causes a measure and / or value (D KL indicative of the similarity or dissimilarity between the first and second output signals to be measured, and first and second output measurements are selected to be those measurements associated with candidate parameters associated with the value or measure indicative of the lowest similarity or highest dissimilarity.
6. The apparatus of claim 1, further comprising an optimizer (560) for iteratively performing the optimization, wherein, the optimizer is configured to randomly generate, at each iteration, a current candidate vector to evaluate whether the current candidate vector performs better than a current best candidate vector, wherein the optimizer is configured to evaluate a target function associated with a similarity or dissimilarity between physical signals associated with the current candidate vector, wherein the optimizer is configured to cause, in case the current candidate vector causes the target function to decrease with respect to the current best candidate vector, the current candidate vector to become a new current best candidate vector.
7. The apparatus of claim 1, configured to: combine the first input signal (502, M0) or a processed version thereof with a delayed and scaled version (5031) of the second input signal (502, M1) in time domain and / or in z-transform or frequency domain; combine the second input signal (502, M1) or a processed version thereof with a delayed and scaled version (5030) of the first input signal (502, M0) in time domain and / or in z-transform or frequency domain.
8. The apparatus of claim 1, wherein the optimization is performed in time domain and / or in z-transform or frequency domain.
9. The apparatus of claim 1, wherein, a delay or fractional delay (d0) applied to the second input signal (502, M1) indicates a relationship and / or a difference or arrival between: a signal from a first source (source0) received by a first microphone (mic0); and a signal from the first source (source0) received by a second microphone (mic1). a delay or fractional delay (d1) applied to the first input signal (502, M0) indicates a relationship and / or a difference or arrival between:
10. The apparatus of claim 1, wherein, a signal from a second source (source1) received by the second microphone (mic1); and a signal from the second source (source1) received by the first microphone (mic0).
11. The apparatus of claim 1, configured to perform the optimization such that different candidate parameters are iteratively selected and processed and a metric is measured for each of the candidate parameters, wherein the metric is a similarity metric or a dissimilarity metric in order to select the first and second input signals by using the candidate parameters associated with the metric indicating the lowest similarity or the highest dissimilarity. the metric (532) is a similarity metric or a dissimilarity metric in order to process and combine the first (502, M0) and second (502, M1) input signals by using the candidate parameters (a0, a1, d0, d1) associated with the metric indicating the lowest similarity or the highest dissimilarity of the output signal. 12. The apparatus of claim 1, configured to perform an optimization (560) such that different candidate parameters (a0, a1, d0, d1) are iteratively selected and processed and a metric (532) is measured for each of the candidate parameters, wherein, 13. The apparatus of claim 11 or 12, wherein, For each iteration, the candidate parameters comprise a candidate delay (d0) to be applied to the second input signal (502, M1), the candidate delay (d0) being associated with a candidate relation and / or a candidate difference or arrival between: a signal from the first source (source0) received by the first microphone (mic0); and a signal from the first source (source0) received by the second microphone (mic1).
14. The apparatus of claim 11, wherein, For each iteration, the candidate parameters comprise a candidate delay (d1) to be applied to the first input signal (502, M0), the candidate delay (d1) being associated with a candidate relation and / or a candidate difference between: a signal from the second source (source1) received by the second microphone (mic1); and a signal from the second source (source1) received by the first microphone (mic0).
15. The apparatus of claim 11, wherein, For each iteration, the candidate parameters comprise a candidate relative attenuation value (564, a0) to be applied to the second input signal (502, M1), the candidate relative attenuation value (564, a0) being indicative of a candidate relation and / or a candidate difference between: an amplitude of a signal received by the first microphone (mic0) from the first source (source0); and an amplitude of a signal received by the second microphone (mic1) from the first source (source0).
16. The apparatus of claim 11, wherein, For each iteration, the candidate parameters (564) comprise a candidate relative attenuation value (a1) to be applied to the first input signal (502, M0), the candidate relative attenuation value (a1) being indicative of a candidate relation and / or a candidate difference between: an amplitude of a signal received by the second microphone (mic1) from the second source (source1); and an amplitude of a signal received by the first microphone (mic0) from the second source (source1).
17. The apparatus of claim 11, configured to vary at least one candidate parameter for different iterations.
18. The apparatus of claim 11, configured to vary at least one candidate parameter for different iterations by randomly selecting at least one step from at least one candidate parameter for a previous iteration to at least one candidate parameter for a subsequent iteration.
19. The apparatus of claim 18, configured to randomly select the at least one step.
20. The apparatus of claim 19, wherein, At least one step is weighted by a pre-selected weight.
21. The apparatus of claim 19 or 20, wherein, At least one step is limited by a pre-selected weight.
22. The apparatus of claim 11, wherein, Candidate parameters (a0, a1, d0, d1) form a vector of candidates, wherein for each iteration the vector of candidates is perturbed by applying a vector of random numbers, the vector of random numbers being multiplied element-wise with or added to elements of the vector of candidates.
23. The apparatus of claim 22, wherein, For each iteration, the vector of candidates is corrected for a step.
24. The apparatus of claim 11, wherein, The number of iterations is limited to a predetermined maximum number.
25. The apparatus of claim 11, wherein, The metric (532) is based on: for each of the first and second signals (M0, M1), a respective element of a set of normalized amplitude values.
26. The apparatus of claim 25, wherein, For at least one of the first and second input signals (M0, M1), the respective element is based on a candidate first or second output signal (S'0, S'1) obtained from the candidate parameters.
27. The apparatus of claim 26, wherein, For at least one of the first and second input signals (M0, M1), the respective element is obtained as a fraction between: a value associated with a candidate first or second output signal (S'0, S'1); and a norm associated with a value of a previously obtained first or second output signal.
28. The apparatus of claim 11, wherein, The metric comprises a logarithm of a quotient formed on the basis of: o a respective element of a first set of normalized amplitude values; and o a respective element of a second set of normalized amplitude values, so as to obtain (530) a value (532) between a signal portion (s0') described by the first set of normalized amplitude values (P0) and a signal portion (s1') described by the first set of normalized amplitude values (P0).
29. The apparatus of claim 11, configured to perform the optimization using a sliding window.
30. The apparatus of claim 1, further configured to transform information associated with the obtained first and second output signals (S'0, S'1) into a frequency domain.
31. The apparatus of claim 1, further configured to encode information associated with the obtained first and second output signals (S'0, S'1).
32. The apparatus of claim 1, further configured to store information associated with the obtained first and second output signals (S'0, S'1).
33. The apparatus of claim 1, further configured to transmit information associated with the obtained first and second output signals (S'0, S'1).
34. An apparatus for a conference call, comprising the apparatus of claim 1 and a device for transmitting information associated with the obtained first and second output signals (S'0, S'1).
35. A binaural system, comprising the apparatus of claim 1 or 34.
36. A method of obtaining a plurality of output signals (504, S'0, S'1) associated with different sound sources (source0, source1) on the basis of a plurality of input signals (502, M0, M1) in which signals from the different sound sources (source0, source1) are combined, the method comprising: combining a first input signal (502, M0) or a processed version thereof with a delayed and scaled version (5031) of a second input signal (502, M1) to obtain a first output signal (504, S'0); combining a second input signal (502, M1) or a processed version thereof with a delayed and scaled version (5030) of a first input signal (502, M0) to obtain a second output signal (504, S'1); determining at least one of: a first scaling value (564, a0) for obtaining a delayed and scaled version (5030) of the first input signal (502, M0); a first delay value (564, d0) for obtaining the delayed and scaled version (5030) of the first input signal (502, M0); a second scaling value (564, a1) for obtaining a delayed and scaled version (5031) of the second input signal (502, M1); and a second delay value (564, d1) for obtaining the delayed and scaled version (5031) of the second input signal, wherein the random direction optimization (560) is such that a candidate parameter forms a candidate vector (564, a0, a1, d0, d1), wherein the candidate vector is iteratively refined by correcting the candidate vector in a random direction, wherein the random direction optimization (560) is such that a measure indicative of a similarity or dissimilarity between the first output signal and the second output signal is measured, and the first output signal and the second output signal are selected as those output signals obtained using the candidate parameters associated with the measure indicative of the lowest similarity or the highest dissimilarity, wherein the measure (532) is processed using a Kullback-Leibler divergence.
37. A non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method of claim 36.