Extraction of an audio object

EP4035154B1Active Publication Date: 2025-06-25LAWO HLDG AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
EP2021701876
Authority / Receiving Office
EP · EP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-14
Filing Date
2021-02-05
Publication Date
2025-06-25
Estimated Expiration
2041-02-05

Smart Images

  • Figure IMGF0001
    Figure IMGF0001
  • Figure IMGF0002
    Figure IMGF0002
  • Figure IMGF0003
    Figure IMGF0003
Patent Text Reader

Abstract

The invention relates to a method for extracting at least one audio object from at least two audio input signals each containing the audio object. According to the invention, the following steps are provided: synchronising the second audio input signal with the first audio input signal obtaining a synchronised second audio input signal; extracting the audio object by applying at least one trained model to the first audio signal and to the synchronised second audio input signal and outputting the audio object. The invention further provides that the method step of synchronising the second audio input signal with the first audio input signal comprises the following method steps: generating audio signals; analytically calculating a correlation between the audio signals; optimising the correlation vector; and determining the synchronised second audio input signal with the aid of the optimised correlation vector. The invention also provides a system having a control unit which is designed to perform the method according to the invention. A computer program containing program code means is also provided, the program being designed to perform the steps of the method according to the invention.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for extracting at least one audio object from at least two audio input signals, each containing the audio object. Furthermore, the invention relates to a system for extracting an audio object and a computer program with program code means.

[0002] For the purposes of the invention, audio objects are audio signals from objects, such as the sound of a soccer ball being kicked, the clapping of an audience, or the speech of a conversation participant. The extraction of the audio object within the meaning of the invention is therefore the separation of the audio object from other, interfering influences, which are referred to below as background noise. For example, when extracting the sound of a shot during a soccer game, the pure shot sound is separated as an audio object from the sounds of the players and the audience, so that the shot sound is ultimately present as a pure audio signal.

[0003] State-of-the-art methods for extracting audio objects are known. A fundamental challenge is that the microphones are typically positioned at different distances from the source of the audio object. Therefore, the audio object is located at different temporal positions relative to the input audio signals, which complicates and slows down the analysis.

[0004] It is known to synchronize the audio input signals in such a way that the audio object is located at the same temporal position as the audio input signals. This is also commonly referred to as runtime compensation. Common methods for this purpose use neural networks. This requires that the neural network be trained for all possible microphone distances from the source of the audio object. However, effective training of the neural network is not feasible, especially with dynamic audio objects, such as in the case of sporting events.

[0005] Furthermore, generic methods are known in which the correlation, for example, their cross-correlation, is analytically calculated to synchronize the audio input signals. While this increases the speed of the process, it impairs the reliability of the subsequent extraction of the audio object, since the correlation is always calculated independently of the type of audio object. However, this often amplifies disruptive effects, particularly noise, for the subsequent extraction of the audio object.

[0006] Well-known generic methods are described in CN 110 534 127 A and in Luo et al: "FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio Processing", AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING WORKSHOP (ASRU), IEEE, December 14, 2019, pages 260-267, XP033718887, DOI: 10.1109 / ASRU46091.2019.9003849 [accessed 2020-02-19].

[0007] It is therefore the object of the invention to eliminate the disadvantages of the prior art mentioned above and, in particular, to improve the reliability of the extraction of the audio object while simultaneously optimizing the speed of the method.

[0008] The object is achieved by a method having the features of claim 1, a system having the features of claim 11 and a computer program having the features of claim 14. Preferred embodiments of the invention are set out in the dependent claims 2-10, 12 and 13.

[0009] The invention is based on the fundamental idea that the analytical calculation of correlation, for example, cross-correlation, improves the quality of the extracted audio object, i.e., the signal separation quality of the method. At the same time, the training of the first and second trained operators creates a possibility of improving the reliability of the subsequent extraction of the audio object with the help of trained components. In this respect, the invention represents a novel method that performs the extraction of the audio object reliably and quickly. This makes the method applicable even with complex microphone geometries, such as large microphone spacing.

[0010] The first trained operator may comprise a particularly trained transformation of the audio input signals into a feature space to simplify the subsequent method steps. The second trained operator may comprise at least one normalization of the correlation vector to improve the accuracy of the calculation of the synchronized second audio input signal. Furthermore, the second trained operator may provide a transformation of the synchronized second audio input signal that is inverse to the transformation of the first trained operator, in particular back to the time period of the audio input signals.

[0011] According to the invention, the second trained operator has an iterative method with a finite number of iteration steps, wherein, in particular, a synchronization vector, preferably an optimized correlation vector, in particular an optimized cross-correlation vector, is determined in each iteration step, which accelerates the method according to the invention. The number of iteration steps of the second trained operator can be user-defined in order to configure the method.

[0012] In each iteration step of the second trained operator, a stretched convolution of the audio signal with at least a portion of the synchronization vector, in particular the optimized correlation vector, preferably occurs. In each iteration step, a normalization of the synchronization vector and / or a stretched convolution of the synchronized audio input signal with the synchronization vector can be performed to improve the signal separation quality of the method.

[0013] In a further embodiment of the invention, the second trained operator determines at least one acoustic model function. Within the meaning of the invention, the acoustic model function corresponds in particular to the relationship between the audio object and the recorded audio input signal. The acoustic model function thus represents, for example, the acoustic properties of the environment, such as acoustic reflections (reverberation), frequency-dependent absorptions, and / or bandpass effects. Furthermore, the acoustic model function includes, in particular, the recording characteristics of at least one microphone. In this respect, the second trained operator can compensate for undesirable acoustic effects on the audio signal, caused, for example, by the environment and / or the recording characteristics of the at least one microphone, as part of the optimization of the correlation vector.In addition to compensating the propagation time, it is also possible to compensate for disturbing acoustic influences, for example caused by the propagation path of the sound, which improves the signal separation quality of the method according to the invention.

[0014] The trained model for extracting the audio object can provide at least one transformation of the first audio input signal and the synchronized second audio input signal, each into a particularly higher-dimensional representation space, which improves the signal separation quality. According to the invention, the representation space has a higher dimensionality than the generally one-dimensional time period of the audio input signals. Since the transformations can be implemented as parts of a neural network, the transformations can be trained specifically with respect to the audio object to be extracted.

[0015] The trained model for extracting the audio object may include applying at least one trained filter mask to the first audio input signal and to the synchronized second audio input signal. The trained filter mask is preferably specifically trained for the audio object.

[0016] The trained model for extracting the audio object may provide at least one transformation of the audio object into the time period of the audio input signals, in particular to undo a previous transformation into the representation space.

[0017] The method steps of synchronizing and / or extracting and / or outputting the audio object are preferably assigned to a single neural network to enable specific training of the neural network with respect to the audio object. The design of a single neural network improves the reliability of the method and its overall signal separation quality.

[0018] Preferably, the neural network is trained with target training data, wherein the target training data comprises audio input signals and corresponding predefined audio objects, with the following training steps: forward feeding the neural network with the target training data while obtaining a determined audio object, determining an error parameter, in particular an error vector between the determined audio object and the predefined audio object and changing parameters of the neural network by backward feeding the neural network with the error parameter, in particular with the error vector, if a quality parameter of the error parameter, in particular of the error vector, exceeds a predefined value.

[0019] The training is focused on the specific audio object; at least two parameters of the trained components of the method according to the invention can be mutually dependent.

[0020] The method is preferably designed to run continuously, which is also referred to as "online operation." According to the invention, audio input signals are continuously read in, in particular without user input, and evaluated to extract audio objects. For example, the audio input signals can each be parts of continuously read audio signals, in particular with a predefined length. This is also referred to as "buffering." Particularly preferably, the method can be designed such that the latency of the method is at most 100 ms, in particular at most 80 ms, and preferably at most 40 ms. According to the invention, latency is the runtime of the method, measured from the reading of the audio input signals to the output of the audio object. Operation of the method in real time is therefore possible.

[0021] The system according to the invention can provide a first microphone for receiving the first audio input signal and a second microphone for receiving the second audio input signal, wherein the microphones can each be connected to the system in such a way that the audio input signals of the microphones can be fed to the control unit of the system. The system can in particular be designed as a component of a mixing console to which the microphones can be connected. Particularly preferably, the system is a mixing console. The connection of the system to the microphones can be wired and / or wireless. The computer program for carrying out the method according to the invention can preferably be executed on a control unit of the system according to the invention.

[0022] Further advantages and features of the invention will become apparent from the claims and the following description, in which embodiments of the invention are explained in detail with reference to the drawings. In the drawings: Fig. 1A system according to the invention in a schematic view; Fig. 2An overview of a method according to the invention in a flowchart with model signals; Fig. 3A flowchart for the method step of synchronizing audio input signals with model signals; Fig. 4A flowchart for an iterative synchronization method; Fig. 5A flowchart for extracting the audio object; and Fig. 6A flowchart for training the method according to the invention.

[0023] Fig. 1 shows an embodiment of a system 10 according to the invention for extracting an audio object 11 in a schematic representation, wherein the system 10 is a mixing console 10a. Audio objects 11 in the sense of the invention are acoustic signals that are associated with an event and / or an object. In the present embodiment of the invention, the audio object 11 is the sound 12 of a fired, Fig. 1 football not shown.

[0024] The noise 12 is recorded by two microphones 13, 14, each generating an audio input signal a1, a2, such that the audio input signals a1, a2 contain the noise 12. Due to the different distances of the microphones 13, 14 from the noise 12, the noise 12 is located at different temporal positions of the audio input signals a1, a2. In addition, the audio input signals a1, a2 differ from one another due to the acoustic properties of the environment and therefore each contain undesirable components, which are caused, for example, by the propagation paths of the sound to the microphones 13, 14, for example in the form of reverberation and / or suppressed frequencies, and which are referred to as noise within the meaning of the invention.According to the invention, a first acoustic model function M1 represents the acoustic influences of the environment and the recording characteristics of the microphone 13 on the recorded audio input signal a1 of the first microphone 13. The audio input signal a1 mathematically corresponds to a convolution of the noise 12 with the first acoustic model function M1. The same applies analogously to a second acoustic model function M2 and to the recorded audio input signal a2 of the second microphone 14.

[0025] The microphones 13, 14 are connected to the mixing console 10a, so that the audio input signals a1, a2 are transmitted to a control unit 15 of the system 10. The control unit 15 evaluates the audio input signals a1, a2 and extracts the sound 12 from the audio input signals a1, a2 using the method according to the invention and outputs it for further use. The control unit 15 for extracting the audio object 11 is a microcontroller and / or a program code block of a corresponding computer program. The control unit 15 comprises a trained neural network that is fed, in particular, forward with audio input signals a1, a2. The neural network is trained to extract the specific audio object 11, in this case the noise 12, from the audio input signals a1, a2 and, in particular, to separate it from noise components of the audio input signals a1, a2.Essentially, the effects of the acoustic model functions M1, M2 on the noise 12 in the audio input signals a1, a2 are compensated.

[0026] Fig. 2 illustrates an embodiment of the method according to the invention in an overview as a flow chart with model audio input signals a1, a2 on which the method is carried out. In a first step V1, the second audio input signal a2 is synchronized with the first audio input signal a1, so that a synchronized second audio input signal a2' is obtained as a result. In the sense of the invention, the synchronized second audio input signal a2' in particular has the noise 12 at essentially the same time position as the first audio input signal a1, which significantly accelerates and simplifies the subsequent method steps. In this respect, the synchronization V1 of the audio input signals a1, a2 corresponds in particular to a compensation of the propagation time differences between the audio input signals a1, a2.

[0027] Subsequently, according to Fig. 2 Extracting V2 of the sound 12 by applying a trained model to the first audio input signal a1 and to the synchronized second audio input signal a2', resulting in the sound 12 as an audio signal. The trained model is assigned to the neural network and, as part of it, is trained to extract the specific audio object 11, here the sound 12. In the subsequent method step, the output V3 of the sound 12 is output as the audio output signal Z.

[0028] The process steps of synchronization (V1), extraction (V2) of the sound (12), and its output (V3) are assigned to a single, trained neural network, so that the process is designed as an end-to-end process. This means that the process is trained as a whole and runs automatically and continuously, with the sound extraction taking place in real time, i.e., with a latency of no more than 40 ms.

[0029] Fig. 3 shows a process sequence V1 of the synchronization of the audio input signals a1, a2 in a flowchart with model audio input signals a1, a2 to illustrate the process steps. In a first process step V4 of the Fig. 3 A first trained operator of the neural network is applied to the audio input signals a1, a2 in order to generate audio signals m1, m2. In one embodiment of the invention, the audio input signals a1, a2 are transformed by the first trained operator of the neural network into a higher-dimensional feature space in the time domain compared to the audio input signals a1, a2 to form the audio signals m1, m2 in order to simplify and accelerate the subsequent calculations. Depending on the type of audio object 11, the audio signals m1, m2 are already processed during the transformation. The transformed audio signals m1, m2 are in Fig. 3 presented as a model.

[0030] In the second process step V5 of the Fig. 3 The analytical calculation of the cross-correlation is carried out as a correlation between the audio signals m1, m2, which is mathematically defined as follows: m 1 ⋆ m 2 t ≙ ∑ n = − ∞ ∞ m 1 n m 2 n + t

[0031] The calculation V5 results in a cross-correlation vector k, which is modeled in Fig. 3 In the third method step V6, the cross-correlation vector k is optimized using a second trained operator of the neural network, whereby the second trained operator is used to calculate the acoustic model function M in order to compensate for its effects on the audio signals m1, m2. The second trained operator thus serves, for example, as an acoustic filter and, in the exemplary embodiment, Fig. 3 In particular, a normalization of the cross-correlation vector k is provided, for example, using a softmax function. The resulting synchronization vector s is modeled in Fig. 3 shown.

[0032] In the fourth step of the Fig. 3 the calculation V7 of the synchronized second audio input signal a2' is carried out by convolving the synchronization vector s with the second audio input signal a2.

[0033] The synchronized second audio input signal a2' is in Fig. 3 represented as a model. In comparison to the original audio input signal a2, it can be seen that in the highly simplified model considered here, the propagation time difference has been compensated for as a temporal offset. The synchronized second audio input signal a2' is then used, as already described, for the extraction V2 of the audio object 11.

[0034] Fig. 4 shows a further embodiment of the synchronization V1 of the audio input signals a1, a2, in which an iterative method is provided to accelerate the calculation, wherein the number of iteration steps I is user-defined. In the first iteration step, the correlation vector between the audio signals m1, m2 is calculated similarly to the method according to Fig. 3 until the calculation V7 of the synchronized audio input signal a2', whereby the synchronization vector si of the current iteration step i is now limited at each iteration step i using the maxpool function as part of the optimization V6. Subsequently, in each iteration step i, the calculation V8 of the iterative audio signal m2 i for iteration step i is performed using a stretched convolution, which is mathematically defined as follows: a 2 ∗ d i s t = ∑ n = − d i d i a 2 d i ⋅ n s n + t

[0035] The factor di corresponds to the degree of restriction of the cross-correlation vector for iteration step i, with summation occurring via the factor di. This process is repeated until the user-specified number of iteration steps I has been performed. Finally, a stretched convolution V9 of the audio signal m2 with the last calculated synchronization vector S i is performed, after which the synchronized second audio signal a2' is calculated and output V7. By calculating the synchronization vector s based on the subrange of the parameters determined in the previous iteration step, the complexity of the calculations is reduced, which accelerates the runtime of the method without compromising its accuracy.

[0036] Fig. 5 shows an embodiment of the extraction V2 of the audio object 11 from the audio input signal a1 and the synchronized second audio input signal a2' in a flowchart. In a first method step V10, the audio input signals a1, a2' are each transformed into a higher-dimensional representation space by applying a first trained model of the neural network in order to simplify the subsequent calculations. For example, the first trained model has a common filter bank with, in particular, a third-octave band filter bank and / or a Mel filter bank, wherein the filter parameters have been optimized by the previous training of the neural network.

[0037] In the second method step V11, the audio object 11 is separated from the audio input signals a1, a2' by applying a second trained model of the neural network to the audio input signals a1, a2'. The parameters of the second trained model were also optimized through the previous training and depend in particular on the first trained model of the previous method step V10. As a result of this method step V11, the audio object 11 is obtained from the audio input signals a1, a2' and is still located in the higher-dimensional representation space.

[0038] In the third process step V12 of the Fig. 5 The separated audio object 11 is transformed into the original, one-dimensional time period of the audio signals a1, a2 by applying a third trained model of the neural network to the audio object 11, whereby the parameters of the third trained model depend on those of the other trained models and were jointly optimized by the previous training. In this respect, the third trained model of the transformation according to the third method step V12 of the Fig. 5 can be seen functionally as a complement to the transformation V10 according to the first trained model. For example, if a one-dimensional convolution is provided in the first trained model of the first process step V10, a transposed one-dimensional convolution is performed in the inverse transformation V12.

[0039] In order for the neural network to reliably extract the audio object 11 from the audio input signals a1, a2, it must be trained before use. This is done, for example, by the training steps V13 to V19 described below, which are described in Fig. 6 are shown in a schematic flow diagram. In the considered embodiments of the method according to the invention, the aforementioned method steps are assigned to a single neural network and are each differentiable, so that with the training method V13 described below, all trained components are trained specifically with regard to the audio object 11.

[0040] Predefined audio objects 16 are generated V14 using predefined algorithms for given audio input signals a1, a2. The predefined audio objects 16 are always of the same type, so that the method is specifically trained with respect to one type of audio object 16. The generated audio input signals a1, a2 run through the inventive method according to Fig. 2 and are fed forward, in particular, by the neural network V15. The audio object 17 thus determined is compared with the predefined audio object 16 in order to determine a mathematical error vector P V16 on this basis. A query V17 then follows as to whether a quality parameter of the error vector P falls below a predefined value and whether the determined audio object 17 was sufficiently well extracted.

[0041] If the quality parameter exceeds the predefined value, the termination criterion is not met and in the next process step V18 the gradient of the error vector P is determined and fed backward through the neural network so that all parameters of the neural network are adjusted. The training process V13 is then repeated with additional data sets until the error vector P reaches a sufficiently good value and the query V17 shows that the termination criterion has been met. The training process V13 is then completed V19 and the process can be applied to real data. Ideally, those audio objects 11 that are also to be determined in the application of the process are used as predefined audio objects 16 in the training phase, for example previously recorded shooting sounds 12 of footballs.

Claims

1. Method for extracting at least one audio object (11) from at least two audio input signals (a1, a2), each of which contains the audio object (11), the method comprising the following steps: - synchronizing (V1) the second audio input signal (a2) with the first audio input signal (a1) to obtain a synchronized second audio input signal (a2'), - extracting (V2) the audio object (11) by applying at least one trained model to the first audio signal (a1) and to the synchronized second audio input signal (a2'), and - outputting (V3) the audio object (11), wherein the method step of synchronizing (V1) the second audio input signal (a2) with the first audio input signal (a1) comprises the following method steps: - generating (V4) audio signals (m1, m2) by applying a first trained operator to the audio input signals (a1, a2), - analytically calculating (V5) a correlation between the audio signals (m1, m2) to obtain a correlation vector (k), - optimizing (V6) the correlation vector (k) using a second trained operator to obtain a synchronization vector (s) and - determining (V7) the synchronized second audio input signal (a2') using the synchronization vector (s), characterized in that the second trained operator has an iterative method with a finite number of iteration steps (I).

2. Method according to claim 1, characterized in that the first trained operator comprises a transformation, in particular a trained transformation, of the audio input signals (a1, a2) into a feature space.

3. Method according to either claim 1 or 2, characterized in that the second trained operator comprises at least one normalization of the correlation vector (k).

4. Method according to any of claims 1 to 3, characterized in that, in each iteration step, a synchronization vector (s) is determined.

5. Method according to claim 4, characterized in that the number of iteration steps (I) of the second trained operator can be defined on the user side.

6. Method according to any of claims 4 to 5, characterized in that, in each iteration step, a normalization of the synchronization vector (s) and / or a stretched convolution of the synchronized audio input signal (a2') with the synchronization vector (s') takes place.

7. Method according to any of claims 1 to 6, characterized in that the second trained operator provides the determination of at least one acoustic model function (M).

8. Method according to any of claims 1 to 7, characterized in that the trained model for extracting (V2) the audio object (11) provides at least one transformation of the first audio input signal (a1) and the synchronized second audio input signal (a2'), each into a representation space, in particular a higher-dimensional representation space and / or in that the trained model for extracting (V2) the audio object (11) provides the application of at least one learned filter mask to the first audio input signal (a1) and to the synchronized second audio input signal (a2'), wherein, in particular, the trained model for extracting (V2) the audio object (11) provides at least one transformation of the audio object (11) into the time period of the audio input signals (a1, a2).

9. Method according to any of claims 1 to 8, characterized in that the method steps of synchronizing (V1) and / or extracting (V2) and / or outputting (V3) the audio object (11) are assigned to a single neural network, wherein, in particular, the neural network is trained with target training data, wherein the target training data comprise audio input signals (a1, a2) and corresponding predefined audio objects (16), the method comprising the following training steps: - forward feeding (V15) the target training data to the neural network to obtain an ascertained audio object (17), - determining (V16) an error vector (P) between the ascertained audio object (17) and the predefined audio object (16), and - changing parameters of the neural network by backward feeding (V18) the error vector (P) to the neural network if a quality parameter of the error vector (P) exceeds a predefined value.

10. Method according to any of claims 1 to 9, characterized in that the method is designed in such a way that it runs continuously and / or in that the audio input signals (a1, a2) are each parts of audio signals (b1, b2), which audio signals are in particular continuously read and have in particular predefined time lengths.

11. System (10) for extracting an audio object (11) from at least two audio input signals (a1, a2), which system has a control unit (15) which is configured to carry out a method according to any of claims 1 to 10.

12. System according to claim 11, characterized in that a first microphone (13) for receiving the first audio input signal (a1) and a second microphone (14) for receiving the second audio input signal (a2) can each be connected to the system (10) in such a way that the audio input signals (a1, a2) of the microphones (13, 14) can be supplied to the control unit (15).

13. System according to either claim 11 or 12, characterized in that the system (10) is configured as a component of a mixing console (10a).

14. Computer program having program code means, which computer program is configured to carry out the steps of a method according to any of claims 1 to 10 when the computer program is run on a computer or a corresponding computing unit, in particular on a control unit (15) of a system (10) according to any of claims 11 to 13.

Citation Information

Patent Citations

  • Microphone array voice enhancement method and device applied to indoor environment

    CN110534127A