Method and apparatus for determining a depth filter

By combining complex time-frequency filters with deep neural networks and training deep filters to process multidimensional tensors, the difficulties of signal extraction and separation in existing technologies are solved, more efficient signal processing effects are achieved, and it is suitable for signal processing tasks in various noisy environments.

CN114041185BActive Publication Date: 2025-09-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080043612.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-16
Filing Date
2020-04-15
Publication Date
2025-09-23
Estimated Expiration
2040-04-15

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively extracting or reconstructing desired signals from mixed signals when dealing with signal extraction and separation in complex noisy environments, especially in the face of destructive interference and signal loss. In particular, the performance of existing methods is limited when extracting speech or biomedical signals in noisy environments.

Method used

Complex time-frequency filters are combined with deep neural networks. Deep filters are estimated by training deep neural networks to process multidimensional tensors to overcome destructive interference and achieve signal extraction, separation and reconstruction.

Benefits of technology

Through the deep filter method, the desired signal can be extracted and separated from the mixed signal more effectively, the signal quality can be improved, and it is suitable for signal processing tasks in various noisy environments, including packet loss concealment and bandwidth extension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114041185B_ABST
    Figure CN114041185B_ABST
Patent Text Reader

Abstract

A method for determining a deep filter, comprising the steps of: receiving a mixture; estimating the deep filter using a deep neural network, wherein the estimation is performed so that the deep filter, when applied to elements of the mixture, obtains an estimate of each element of a desired representation; wherein the deep filter having at least one dimension comprises a tensor having elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a method and apparatus for determining a depth filter. Further embodiments relate to the use of the method for signal extraction, signal separation or signal reconstruction. Background Art

[0002] When a signal is captured by a sensor, it usually contains desired and undesired components. Consider speech (desired) in a noisy environment with additional interfering speakers or directional noise sources (undesired). The desired speech is extracted from the mixture to obtain a high-quality noise-free recording and can benefit the perceived speech quality, for example in teleconferencing systems or mobile communications. Considering different scenarios in which biomedical signals such as electrocardiograms, electromyograms or electroencephalograms are captured by sensors, interference or noise must also be eliminated to achieve an optimal interpretation and further processing of the captured signal, for example by a doctor. In general, in a variety of different scenarios, it is desirable to extract the desired signal from a mixture or to separate multiple desired signals in a mixture.

[0003] Beyond extraction and separation, there are also scenarios where parts of the captured signal are no longer accessible. Consider transmission scenarios where some packets are lost, or audio recordings where room acoustics cause spatial comb filters and cancel / corrupt specific frequencies. Assuming that the remaining portion of the signal contains information about the content of the lost part, reconstructing the lost signal portion is also highly desirable in a variety of different scenarios.

[0004] Current signal extraction and separation methods are discussed below:

[0005] Given adequate estimates of the statistics of the desired and undesired signals, traditional methods, such as Wiener filtering, apply real-valued gains to the complex mixture short-time Fourier transform (STFT) representation to extract the desired signal from the mixture [e.g.,

[01] ,

[02] ].

[0006] Another possibility is to estimate from statistics a complex-valued multidimensional filter in the STFT domain for each mixed time-frequency bin and apply it to perform the extraction. For the separation scenario, each desired signal requires its own filter

[02] .

[0007] Statistics-based methods perform well given stationary signals; however, given highly non-stationary signals, statistical estimation is often challenging.

[0008] Another approach is to use non-negative matrix factorization (NMF). It learns in an unsupervised manner from the provided training data basis vectors of the data, which can be identified during testing [e.g.,

[03] ,

[04] ]. Given that speech must be separated from white noise, NMF learns the most prominent basis vectors in the training examples. Since white noise is not temporally correlated, these vectors belong to speech. During testing, it can be determined whether one of the basis vectors is currently active to perform extraction.

[0009] Speech signals from different speakers are very different, and approximating all possible speech signals with a limited number of basis vectors cannot satisfy this high variance of the expected data. In addition, if the noise is highly unstable and unknown during training, unlike white noise, the basis vectors may overlap the noise segments, which will degrade the extraction performance.

[0010] In recent years, time-frequency masking techniques, especially those based on deep learning, have shown significant improvements in performance [e.g.,

[05] ]. Given labeled training data, a deep neural network (DNN) is trained to estimate a time-frequency mask. This mask is applied element-wise to the complex mixture STFT to perform signal extraction or, in the case of multiple masked signal separation, signal extraction. The mask elements can be binary given a time-frequency grid dominated by only a single source [e.g.,

[06] ]. The mask elements can also be real-valued ratios [e.g.,

[07] ] or complex-valued ratios [e.g.,

[08] ] given multiple active sources per time-frequency grid.

[0011] This extraction is done by Figure 1 Shown. Figure 1 Shows multiple persons x,y These cells are the input STFT, where the region marked by A in the input STFT is fed to the DNN to estimate the gain for each time-frequency cell therein. This gain is applied element-wise to the complex input STFT (see the cells marked by x in the input and extraction diagrams). This is done to estimate the corresponding desired component.

[0012] Given a mixed time-frequency grid of zeros due to destructive interference from the desired and undesired signals, masks cannot reconstruct the desired signal by applying only gain to this grid, as the corresponding mask value does not exist. Even if the mixed time-frequency grid is close to zero due to destructive interference from the desired and undesired signals, masks are generally unable to fully reconstruct the desired signal by applying only gain to this grid, as the corresponding masks are typically limited in magnitude, which limits their performance, given the destructive interference in a particular time-frequency grid. Furthermore, given partial signal loss, masks cannot reconstruct these portions, as they only apply gain to the time-frequency grid to estimate the desired signal.

[0013] Therefore, an improved method is needed. Summary of the Invention

[0014] It is an object of the present invention to provide an improved method for signal extraction, separation and reconstruction.

[0015] This object is solved by the subject-matter of the independent claims.

[0016] Embodiments of the present invention provide a method for determining a deep filter having at least one dimension. The method comprises the steps of receiving a mixture and estimating a deep filter using a deep neural network, wherein the estimation is performed such that, when the deep filter is applied to elements of the mixture, an estimate of each element of a desired representation is obtained. Here, the deep filter having at least one dimension comprises a tensor having elements.

[0017] The present invention is based on the discovery that combining the concept of complex time-frequency filters from the statistical methods section with deep neural networks allows for extracting / separating / reconstructing expected values ​​from multidimensional tensors (assuming that the multidimensional tensor is the input representation). This general framework is called deep filters and is based on distorted / noisy input signals processed using neural networks (which can be trained using cost functions and training data). For example, the tensor can be a one-dimensional or two-dimensional complex STFT or an STFT with an additional sensor dimension, but is not limited to these scenarios. In this paper, deep neural networks are directly used to estimate one-dimensional or even multi-dimensional (complex) deep filters for each equation tensor element (A). These filters are applied to limited regions of the degraded tensor to obtain an estimate of the expected value in the enhanced tensor. In this way, the problem of masks that have destructive interference due to their bounded values ​​can be overcome by merging several tensor values ​​for estimation. Due to the use of DNNs, the statistical estimation of time-frequency filters can also be overcome.

[0018] According to an embodiment, the mixture may include a real-valued or complex-valued time-frequency representation (such as a short-time Fourier transform) or a characteristic representation thereof. In this document, the desired representation also includes the desired real-valued or complex-valued time-frequency representation or its characteristic representation. According to an embodiment, it may turn out that the depth filter also includes a real-valued or complex-valued time-frequency filter. In this case, one dimension of the depth filter can be selected to be described in the short-time Fourier transform domain.

[0019] Furthermore, at least one dimension may be selected from the group consisting of a time dimension, a frequency dimension, or a sensor signal dimension. According to a further embodiment, the estimation is performed for each element of the mixture, or a predetermined portion of an element of the mixture, or a predetermined portion of a tensor element of the mixture. According to an embodiment, this estimation may be performed for one or more, for example, at least two, sources.

[0020] Regarding the definition of filters, it should be noted that, according to embodiments, the method may include the step of defining a filter structure having filter variables for a deep filter having at least one dimension. This step may be associated with embodiments in which the deep neural network includes multiple output parameters, where the number of output parameters may be equal to the number of filter values ​​of the filter function of the deep filter. Note that the number of trainable parameters is typically much larger, where defining the number of outputs equal to the number of real plus imaginary filter components is beneficial. According to embodiments, the deep neural network includes a batch normalization layer, a bidirectional long short-term memory layer, a feedforward output layer, a feedforward output layer with a hyperbolic tangent activation, and / or one or more additional layers. As described above, this deep neural network can be trained. Therefore, according to embodiments, the method includes the step of training the deep neural network. This step may be performed using a training sub-step using the mean squared error (MSE) between the ground truth and a desired representation, as well as an estimate of the desired representation. Note that an exemplary method of the training process is to minimize the mean squared error during DNN training. Alternatively, the deep neural network can be trained by reducing the reconstruction error between the desired representation and the estimate of the desired representation. According to further embodiments, training is performed using magnitude reconstruction.

[0021] According to an embodiment, the estimation may be performed by using the following formula

[0022]

[0023] Where 2·L+1 is the filter dimension in the time frame direction, 2·I+1 is the filter dimension in the frequency direction, is a complex conjugate two-dimensional filter. For the sake of completeness, it should be noted that the above formula Indicates the action that should be performed in the Apply step.

[0024] Starting from this formula, training can be performed using the following formula,

[0025]

[0026] where X d (n,k) is the expected representation, and is the expected representation of the estimate, or

[0027] Use the following formula:

[0028]

[0029] where X d (n,k) is the expected representation, is the expected representation of the estimate.

[0030] According to an embodiment, the elements of the depth filter are bounded in magnitude or are bounded in magnitude by using the following formula,

[0031]

[0032] in is a complex conjugate two-dimensional filter. Note that in the preferred embodiment, the bound is due to the hyperbolic tangent activation function of the DNN output layer.

[0033] Another embodiment provides a filtering method. This method includes the essential steps and optional steps of the above-described method for determining a depth filter, and a step of applying the depth filter to the mixture. It should be noted that, according to an embodiment, the applying step is performed by element-wise multiplication and successive summation to obtain an estimate of the desired representation.

[0034] According to a further embodiment, the filtering method can be used for signal extraction and / or for signal separation of at least two sources. Another application according to a further embodiment is that the method can be used for signal reconstruction. Typical signal reconstruction applications are packet loss concealment and bandwidth extension.

[0035] It should be noted that the method for filtering, as well as the method for signal extraction / signal separation and signal reconstruction, can be performed by a computer. This applies to the method for determining a depth filter having at least one dimension. This means that a further embodiment provides a computer program having program code for performing one of the above methods when executed on a computer.

[0036] Another embodiment provides an apparatus for determining a depth filter. The apparatus includes an input for receiving a mixture;

[0037] A deep neural network for estimating a deep filter that, when applied to a mixture of elements, can obtain an estimate of the corresponding element of the desired representation. Here, the filter includes a tensor of at least one dimension (having elements).

[0038] According to another embodiment, a device capable of filtering a mixture is provided. The device comprises a depth filter as defined above applied to the mixture. The device can be enhanced to enable signal extraction / signal separation / signal reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The embodiments of the present invention will be discussed below with reference to the accompanying drawings, in which

[0040] Figure 1Schematically shown are diagrams representing a mixture as input (frequency-time diagram) and diagrams representing extraction to illustrate the principle of generating / determining a filter according to conventional methods;

[0041] Figure 2a Schematically shows an input diagram (frequency-time diagram) and an extraction diagram (frequency-time diagram) for explaining the principle of an estimation filter according to an embodiment of the present invention;

[0042] Figure 2b A schematic flow chart of a method for determining a depth filter according to an embodiment is shown;

[0043] Figure 3 shows a schematic block diagram of a DNN architecture according to an embodiment;

[0044] Figure 4 shows a schematic block diagram of a DNN architecture according to a further embodiment;

[0045] Figure 5 ab show two graphs showing the MSE results of two tests for illustrating the advantages of the embodiment;

[0046] Figure 6 a-6c schematically show excerpts of logarithmic magnitude STFT spectra for illustrating the principles and advantages of embodiments of the present invention. DETAILED DESCRIPTION

[0047] Hereinafter, embodiments of the present invention will be discussed with reference to the accompanying drawings, wherein the same reference numerals are used for elements / objects having the same or similar functions so that their descriptions are mutually applicable and interchangeable.

[0048] Figure 2a Two frequency-time diagrams are shown, wherein the left frequency-time diagram marked by reference numeral 10 represents the mixture received as input. Here, the mixture is an STFT (Short Time Fourier Transform) with a plurality of bit grids s x,y Some of the bits marked by reference numeral 10a are used as input for the estimation filter, which is Figure 2a and 2b The purpose of the method 100 is described in the background.

[0049] like Figure 2b As shown, the method 100 includes two basic steps 110 and 120. The basic step 110 receives a mixture 110, such as Figure 2a As shown in the left picture.

[0050] In the next step 120, a depth filter is estimated. This step 120 is illustrated by arrow 12, which maps to the marker grid 10x used as the extracted right-hand frequency-time plot. The estimated filter is visualized by the cross 10x and is estimated so that the depth filter, when applied to the mixture elements, yields an estimate of the corresponding element of the desired representation 11 (see abstract diagram). In other words, this means that the filter can be applied to a limited region of the complex input STFT to estimate the corresponding desired component (see extracted diagram).

[0051] Here, DNN is used to represent each degenerate tensor element s x,y Estimate at least one-dimensional or preferably multi-dimensional (complex) depth filters, as shown in 10x. The filter 10x (for degenerate tensor elements) is applied to the degenerate tensor s x,y A bounded region 10a of is defined to obtain an estimate of the expected value in the augmented tensor. In this way, the problem of destructive interference caused by the bounded value of the mask can be overcome by merging several tensor values ​​for the estimate. Note that the mask is bounded because the DNN output is in a finite range, typically (0,1). From a theoretical point of view, the range (0,∞) would be the preferred variant for performing perfect reconstruction, where it has been demonstrated in practice that the above-mentioned limited range is sufficient. Thanks to this approach, the statistical estimation of the time-frequency filter can be overcome by using a DNN.

[0052] about Figure 2a In the example shown, it should be noted that a square filter is used herein, wherein the filter 10 is not limited to this shape. It should also be noted that the filter 10x has two dimensions, namely the frequency dimension and the time dimension, wherein according to another embodiment, the filter 10x may have only one dimension, namely the frequency dimension or the time dimension or another (not shown) dimension. In addition, it should be noted that the filter 10a has more than two dimensions shown, i.e., it can be implemented as a multidimensional filter. Although the filter 10x has been described as a 2D composite STFT filter, another possible option is to implement the filter as an STFT with an additional sensor size, i.e., not necessarily a composite filter. Alternatives are real-valued filters or quaternion value filters. These filters can also have at least one or more dimensions to form a multidimensional depth filter.

[0053] Multidimensional filters offer versatile solutions for a variety of tasks (signal separation, signal reconstruction, signal extraction, noise reduction, bandwidth extension, etc.). They are able to perform better signal extraction and separation than time-frequency masks (the current technique). Because they reduce destructive interference, they can be used for packet loss concealment or bandwidth extension, which are similar problems to destructive interference and therefore cannot be solved by time-frequency masks. In addition, they can be used to declip signals.

[0054] Deep filters can be specified in different dimensions, such as time, frequency, or sensor, which makes them very flexible and applicable to a variety of different tasks.

[0055] Compared to conventional prior art methods, signal extraction from a single-channel mixture with an additive undesired signal, most commonly performed using a time / frequency (TF) mask, clearly demonstrates that the complex TF filter estimated using a DNN is estimated for each mixture TF grid, mapping the corresponding STFT region in the mixture to the desired TF grid to account for destructive interference in the mixture TF grid. As discussed above, the DNN can be optimized by minimizing the error between the extracted and ground-truth desired signals, allowing training without specifying the ground-truth TF filter, but rather learning the filter by minimizing the error. For the sake of completeness, it should be noted that conventional methods for extracting additive undesired signals from single-channel mixtures most commonly use time-frequency (TF) masks. Typically, a deep neural network (DNN) is used to estimate the mask and apply it element-wise to the complex mixture short-time Fourier transform (STFT) representation to perform extraction. The ideal mask amplitude is zero for individual undesired signals in the TF grid, but infinite for the total destructive interference. Typically, the mask has an upper bound to provide a well-defined DNN output at the expense of limited extraction capability.

[0056] The following will refer to Figure 3 The filter design process is discussed in more detail.

[0057] Figure 3 An example DNN architecture is shown for mapping the real and imaginary parts of the input STFT 10 to filters 10x using a DNN 20 (see Figure 3 a). According to Figure 3 In the embodiment shown in b, the DNN architecture may include multiple layers, thereby using three bidirectional long short-term memory layers BLTSMS (or three long short-term memory layers) LSTMS (both plus a feed-forward layer with hyperbolic tangent activation for activating the real and imaginary values ​​of the deep filters) to perform their mapping. Note that BLSTMS has LSTM paths in both the time direction and the reverse time direction.

[0058] The first step is to define a filter structure specific to the problem. Figure 2b ), this optional step is marked by reference numeral 105. The design of this structure is a compromise between computational complexity (i.e. the more filter values, the more computations are required, while the performance given by too few filter values, e.g. destructive interference or data loss, can come into play again, thus giving a reconstruction bound).

[0059] The deep filter 10x is obtained by providing the mixture 10 or a feature representation thereof to the DNN 20. For example, the feature representation may be the real and imaginary parts of a complex mixture STFT as input 10.

[0060] As shown above, the DNN architecture can consist of, for example, a batch normalization layer, a (bidirectional) long short-term memory layer (BLSTM), and a feed-forward output layer with, for example, a hyperbolic tangent activation. The hyperbolic tangent activation results in a DNN output layer in [-1, 1]. A specific example is given in the Appendix. If an LSTM is used instead of a BLSTM, online separation / reconstruction can be performed because the backward path in time is avoided in the DNN structure. Of course, additional or alternative layers can be used within the DNN architecture 10.

[0061] According to a further embodiment, the DNN can be trained using the mean squared error between the true value and the estimated signal given by applying the filter to the mixture. Figure 2 shows the application of an example filter estimated by the DNN. The red cross in the input marks the STFT bit grid for which the complex filter value has been estimated to estimate the corresponding STFT bit grid in the extraction (marked by the red cross). There is one filter estimate for each value in the extraction STFT. Assuming that there are N desired sources separated in the input STFT, the extraction process is performed separately for each of them. The filter must be estimated for each source, for example with Figure 4 The architecture shown in .

[0062] Figure 4 An example DNN architecture is shown that maps the real and imaginary values ​​of the input STFT 10 to a plurality of filters 10x1 to 10xn. Each filter 10x1 to 10xn is designed for a different desired source. This mapping is performed using a DNN 20, as described with respect to Figure 3 discussed.

[0063] According to the embodiment, the estimated / determined depth filter can be used in different application scenarios. The embodiment provides a method for signal extraction and separation by using the depth filter determined according to the above principle.

[0064] When one or more desired signals must be extracted from a mixture STFT, a possible filter form is a 2-dimensional rectangular filter per STFT bit grid for each desired source to perform the separation / extraction of the desired signal. Such a deep filter is Figure 2a As shown in .

[0065] According to a further embodiment, a deep filter may be used for signal reconstruction. If the STFT mixture is degraded by pre-filtering (e.g., a notch filter), clipping artifacts or parts of the desired signal are lost (e.g., due to packets lost during transmission or narrowband transmission [e.g., [9]]).

[0066] In the above cases, the desired signal must be reconstructed using time and / or frequency information.

[0067] The considered scenario has solved the reconstruction problem, where the STFT bit grid is lost in the time or frequency dimension. In the context of bandwidth extension (for example, in the case of narrowband transmission), a specific STFT region (for example, the upper frequency limit) is lost. Based on the prior knowledge about the non-degenerate STFT bit grid, the number of filters can be reduced to the number of degenerate STFT bit grids (i.e., the upper frequency limit is lost). We can retain the rectangular filter structure but apply depthwise filters to the given lower frequencies to perform bandwidth extension.

[0068] The above examples / implementations describe deep filters for signal extraction using complex time-frequency filters. In the following method, methods with complex and real-valued TF masks are compared by separating speech from a variety of different sound and noise categories in the Google AudioSet corpus. In this paper, the mixture STFT is processed with a notch filter and zeroed integer time frames to demonstrate the reconstruction capabilities of the method. The proposed method outperforms the baseline, especially when applying the notch filter and zeroing the time frame.

[0069] Real-world signals are often corrupted by unwanted noise sources or interferers, such as white self-noise from microphones, background sounds like crowd noise or traffic, and impulsive sounds like clapping. Preprocessing, such as notch filtering or specific room acoustics leading to spatial comb filters, can also lead to degradation of the recorded signal. When high-quality signals are required, it is highly desirable to extract and / or reconstruct the desired signal from this mixture. Possible applications are, for example, enhancing recorded speech signals, separating different sources from each other, or concealing packet loss. Signal extraction methods can be broadly categorized into single-channel and multi-channel methods. In this paper, we focus on single-channel methods and address the problem of extracting the desired signal from a mixture of desired and undesired signals.

[0070] Common approaches perform this extraction in the short-time Fourier transform (STFT) domain, where either the desired spectral magnitudes (e.g., [1]) or a time-frequency (TF) mask is estimated and then element-wise applied to a complex mixture STFT to perform the extraction. For performance reasons, estimating the TF mask is often preferred over directly estimating the spectral magnitudes [2]. Typically, the TF mask is estimated from a mixture representation of a deep neural network (DNN) (e.g., [2]–[9]), where the output layer typically directly produces the STFT mask. There are two common approaches to training such DNNs. First, the true mask is bounded and the DNN learns the mixture to perform the mask mapping by minimizing the error function between the true mask and the estimated mask (e.g., [3], [5]). In the second approach, the DNN learns the mapping by directly minimizing the error function between the estimated signal and the desired signal (e.g., [8],

[10] ,

[11] ). Erdogan et al.

[12] showed that direct optimization is equivalent to mask optimization weighted by the squared mixture magnitudes. As a result, the impact of high-energy TF bins on the loss increases, while the impact of low-energy bins decreases. Furthermore, there is no need to define a truth mask since it is implicit in the truth expectation signal.

[0071] Different types of TF masks have been proposed for different extraction tasks. Given a mixture in the STFT domain, where the signal in each TF bin belongs only to the desired signal or the undesired signal, extraction can be performed using binary masks

[13] , which have been used, for example, in [5], [7]. Given a mixture in the STFT domain, where multiple sources are active in the same TF bin, a ratio mask (RM)

[14] or a complex ratio mask (cRM)

[15] can be applied. Both assign a gain to each mixture TF bin to estimate the desired spectrum. The real-valued gain of the RM performs a TF bin-level amplitude correction from the mixture to the desired spectrum. In this case, the estimated phase is equal to the mixture phase. The cRM applies complex rather than real gains and additionally performs a phase correction. Speaker separation, dereverberation, and denoising have been achieved using RMs (e.g., [6], [8],

[10] ,

[11] ,

[16] ) and cRMs (e.g., [3], [4]). Ideally, the magnitudes of RM and cRM are zero if only the undesired signal is active in the TF bit grid, and infinite if the desired and undesired signals destructively overlap in a certain TF bit grid. Outputs approaching infinity cannot be estimated using DNNs. To obtain well-defined DNN outputs, a compressed mask can be estimated using DNNs (e.g., [4]) and extraction can be performed after decompression to obtain high-magnitude mask values. However, weak noise on the DNN output can cause large changes in the estimated mask, resulting in large errors. In addition, when the desired and undesired signals in the TF bit grid add up to zero, the compressed mask cannot be reconstructed from zero by multiplication. Typically, the case of destructive interference is ignored (e.g., [6],

[11] ,

[17] ) and the mask value is estimated to be bounded by 1, because higher values ​​also carry the risk of noise amplification. In addition to masks, complex-valued TF filters (e.g.,

[18] ) have also been used for signal extraction. Current TF filter methods often incorporate a statistics estimation step (e.g.,

[18]

[21] ), which can be crucial given the large number of unknown interfering signals with rapidly changing statistics present in real-world scenarios.

[0072] In this paper, we propose to use a DNN to estimate the complex-valued TF filter for each TF-cell in the STFT domain to solve the problem of extraction of highly non-stationary signals with unknown statistics. The filter is applied element-wise to a defined region in the STFT of the corresponding mixture. The results are summed to obtain an estimate of the desired signal in the corresponding TF cell. The individual complex filter values ​​are bounded in amplitude to provide a well-defined DNN output. Each estimated TF cell is a complex weighted sum of the TF cell regions in the complex mixture. This allows to solve the case of destructive interference in a single TF cell without the noise sensitivity of mask compression. It also allows to reconstruct a TF cell that is zero by considering neighboring TF cells with non-zero amplitudes. The combination of DNN and TF filter alleviates the shortcomings of TF mask and existing TF filter methods.

[0073] This paper is organized as follows. In Section 2, we present the signal extraction process using TF masks, followed by a description of our proposed method in Section 3. Section 4 contains the datasets we used, and Section 5 contains experimental results to validate our theoretical ideas.

[0074] Starting from this extraction, an STFT mask-based extraction is performed. The extraction process using the TF mask is described, and the implementation details of the mask used as a baseline in the performance evaluation are provided.

[0075] A. Goal

[0076] We define the complex single-channel spectrum of the mixture as X(n,k) in the STFT domain and the desired signal as X d (n,k), the undesired signal is limited to X u (n,k), where n is the time frame and k is the frequency index. We consider the mixture X(n,k) to be a superposition

[0077] X(n, k) = X u (n, k)+X d (n, k)· (1)

[0078] Our goal is to obtain X(n,k) by applying the mask to X(n,k) as a superposition d Estimate of (n,k)

[0079]

[0080] in is the estimated expected signal, is the estimated TF mask. For a binary mask, is∈{0,1}, for is ∈[0,b] and the upper bound is for is ∈[0,b] and is ∈ C. The upper limit b is usually 1 or close to 1. The binary mask classifies the TF bit grid, RMs performs amplitude correction, and cRMs additionally performs phase correction

[0081] From X(n,k) to In this case, solving the extraction problem is equivalent to solving the mask estimation problem.

[0082] Typically, TF masks are estimated using a DNN that is optimized to estimate pre-defined ground-truth TF masks for all N·K TF bins, where N is the total number of time frames and K is the number of frequency bins per time frame.

[0083]

[0084] Use the true value mask M(n,k), or reduce the reconstruction

[0085] X d (n,k) and

[0086]

[0087] or amplitude reconstruction

[0088]

[0089] Optimizing the reconstruction error is equivalent to performing a weighted optimization on the mask, reducing the influence of low-energy TF cells and increasing the influence of high-energy TF cells on the loss

[12] . For the destructive interference in (1), the well-known triangle inequality given by

[0090] |X d (n, k)+X u (n, k)|<|X d (n, k)|<|X d (n, k)|+|X u (n, k)|, (6)

[0091] It holds true, requiring 1<|M(n,k)|≤∞. Therefore, the global optimum cannot be reached above the upper limit of the mask b.

[0092] B. Implementation

[0093] For mask estimation, we use a DNN with a batch norm layer, followed by three bidirectional long short-term memory (BLSTM) layers

[22] , each with 1200 neurons, and a feed-forward output layer with hyperbolic tangent activation, producing an output O with dimension ((N, K, 2)) representing the imaginary and real outputs of each TF bit grid ∈ [-1, 1].

[0094] For mask estimation, we design the models to have the same number of trainable parameters and the same maximum We used a real-valued DNN with the stacked imaginary and real parts of X as input and two outputs, and each TF bit grid is limited to O r and O i These can be interpreted as imaginary and real mask components. For RM estimation, we compute get for and The magnitude is between 1 and √2, where for O i (n, k) reaches 1. This setting is for |O r (n, k)|=|O i (n, k)|=1 produces the maximum cRM pure real or imaginary mask value of phase correlation and √2, resulting in the amplification disadvantage of cRM compared with RM. We trained two DNNs to estimate RM optimized using (5) and cRM optimized using (4). We calculated X(n, k) and X(n, k) in (2) for cRM by the following formula Complex multiplication of

[0095] Re{X d}=Re{M}·Re{X}-Im{M}·Im{X}, (7)

[0096]

[0097] Note that (n, k) is omitted for brevity. We trained for 100 epochs using the Adam

[23] optimizer, information loss

[24] of 0.4 in BLSTM, a batch size of 64, an initial learning rate of 1e-4, multiplied by 0.9, and no decrease in validation loss after each epoch.

[0098] In the following, we will discuss the proposed improved method for STFT filter-based extraction. In this paper, we will specifically show how to estimate x using STFT domain filters instead of TF masks. d This filter is called a depth filter (DF).

[0099] A. Goal

[0100] We propose to obtain from X by applying complex filters

[0101]

[0102] Where 2·L+1 is the filter dimension in the time frame direction, and 2·I+1 is in the frequency direction, is a two-dimensional filter that is the complex conjugate of the TF bit grid (n,k). Note that without loss of generality, we use square filters in (9) just for simplicity of representation. The filter values ​​are like mask values ​​with bounded magnitudes to provide well-defined DNN outputs.

[0103]

[0104] The DNN is optimized according to (4), which allows training without having to define the true GTF and directly optimizing the reconstruction mean squared error (MSE). The decision of the GTF is crucial because there are usually countless combinations of different filter values ​​that will lead to the same extraction result. If a GTF is randomly selected for the TF bit grid from an infinite set of GTFs, training will fail because there will be no consistency between the selected filters. We can interpret this situation as a process that is partially observable to the GTF designer but fully observable to the DNN. Based on the properties of the input data, the DNN can accurately decide which filter to use without ambiguity. The GTF designer has an infinite set of possible GTFs but cannot interpret the input data to decide which GTF to use so that the current DNN update is consistent with the previous update. By using (4) for training, we avoid the problem of GTF selection.

[0105] B. Implementation

[0106] We use the same DNN as proposed in Section II-B, changing the output shape to (N, K, 2, 2·L+1, 2·I+1), where the last two entries are the filter dimensions. The complex multiplication in (9) is performed as in (7) and (8). In our experiments, we set L = 2 and I = 1 to obtain the maximum value of the filter |H for the dimension (5, 3). n,k (l,i)| is phase-dependent Similar to the CRM in Subsection II-B, the output layer activation is used. As in all |H n,k (l,i)| can be at least 1. DNN can theoretically optimize (4) to its global optimum of zero if

[0107]

[0108] in is the maximum amplitude that can be achieved by all filter values, and in our setting c = 1. Therefore, to resolve destructive interference, the sum of all mixture amplitudes considered by the filter weighted with c must be at least equal to the desired TF bin amplitude. Since the filter exceeds the spectrum of the TF bin at the edges, we pad the spectrum with L zeros in the time axis and with I in the frequency axis.

[0109] 4. Dataset

[0110] We use AudioSet

[25] as the interference source (no speech samples) and LIBRI

[26] as the expected speech data corpus. All data are downsampled to 8kHz sampling frequency and 5 seconds in duration. For STFT, we set the hop size to 10ms, the frame length to 32ms, and use the Hann window. Therefore, in our test, K = 129 and N = 501.

[0111] We degraded the desired speech samples by adding white noise, interference from AudioSet, notch filtering, and random timeframe zeroing (T-kill). Each degradation was applied to the sample with 50% probability. For AudioSet interference, we randomly selected 5 seconds of AudioSet and desired speech from LIBRI to compute a training sample. Speech and interference were mixed with a piecewise signal-to-noise ratio (SNR) ∈ [0, 6] dB, and speech and white noise with SNR ∈ [20, 30] dB. For notch filtering, we randomly selected center frequencies with quality factors ∈ [10, 40]. When applying T-kill, each timeframe was zeroed with 10% probability. We generated 100,000 training, 5,000 validation, and 50,000 test samples using the corresponding LIBRI set and the aforementioned degradations. To avoid overfitting, training, validation, and test samples were created from different speech and interference samples from AudioSet and LIBRI. We divided the test samples into three subsets: Test 1, Test 2, and Test 3. In Test 1, the speech was degraded solely by the interference of AudioSet. In Test 2, the speech was degraded solely by notch filtering and T-kill. In Test 3, the speech was degraded by interference, notch filtering, and T-kill simultaneously. All subsets include samples with and without white noise.

[0112] D. Performance Evaluation

[0113] For performance evaluation, we used signal-to-distortion ratio (SDR), signal-to-distortion ratio (SAR), signal-to-interference ratio (SIR)

[27] , reconstruction MSE (see (4)), short-term objective intelligibility (STOI)

[28] ,

[29] and the test dataset.

[0114] First, we tested how clean speech degrades when processed. The MSEs after RM, cRM, and DF were applied were -33.5, -30.7, and -30.2 dB, respectively. The errors are very small, and we assume that they are caused by noise on the DNN output. RMs produce the smallest MSE because noise on the DNN output only affects the amplitude, then cRMs are affected as both phase and amplitude are affected, and finally, DF introduces the highest MSE. In informal listening tests, no difference was perceived. Table I shows the average results for Tests 1-3. In Test 1, DF, cRM, and RM were shown to generalize well to unseen perturbations. Using cRM instead of RM for processing did not result in performance improvement.

[0115] Table I: Average results SDR, SIR, SAR, MSE (dB), RM, cRM and DF for test sample degradation with AudioSet interference in test 1, with notch filter and time frame zeroing (T-kill) in test 2, and the combination in test 3; the non-proposed MSE for tests 1, 2, and 3 are 1.60, -7.80, 1.12, respectively, and STOI are 0.81, 0.89, 0.76, respectively

[0116]

[0117] In addition to amplitude correction, phase is also performed. This is probably due to the amplification disadvantage of cRM compared to RM caused by the adopted DNN architecture described in subsection II-B. For the metric STOI, DF and RM perform comparable, while for other metrics, DF performs better and achieves a further improvement of 0.61dB in SDR. The box plots of the MSE results are shown in Figure 5 We assume that this is due to the superior reconstruction capability of DF relative to destructive interference. In Test 2, DF clearly outperforms cRM and RM because the test conditions provide a scenario comparable to destructive interference. Figure 6 The logarithmic magnitude spectra of clean speech, degraded speech after zeroing every five time frames and frequency axes and enhancing with DF are depicted. Unlike the random time frame zeroing in the dataset, Figure 6 The degradation in is for illustration purposes only. Traces of the mesh are still visible in the low-energy spectral region, but not in the high-energy spectral region, as the loss in (4) focuses on. In Test 3, DFs perform best as they are able to resolve all degradation issues, while RM and cRM cannot. The baseline cRM and RM perform the same.

[0118] The conclusions are as follows:

[0119] We extend the concept of time-frequency masks for signal extraction to complex filters to increase interference reduction and reduce signal distortion, and to account for destructive interference from both desired and undesired signals. We propose using a deep neural network to estimate the filters, which is trained by minimizing the mean square error (MSE) between the desired and estimated signals. This avoids defining the true filters for training, which is crucial because consistently defining the filters is necessary for training the network given an infinite number of possibilities. Our filter and mask approach is able to perform speech extraction given unknown interfering signals from an audio set, demonstrating its versatility and introducing only minimal error when processing clean speech. Our approach outperforms both complex ratio masks and ratio mask baselines in all but one metric, where performance is comparable. In addition to interference reduction, we also test whether simulated data loss, such as by time frame zeroing or filtering with a notch filter, can be addressed, and show that only our proposed approach can reconstruct the desired signal. Thus, signal extraction and / or reconstruction appears feasible under highly adverse conditions, given packet loss or unknown interference, using deep filters.

[0120] As mentioned above, the above methods can be executed by a computer, that is, the embodiment refers to a computer program for executing one of the above methods. Similarly, the method can be executed by using an apparatus.

[0121] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or items or features of corresponding apparatus. Some or all method steps can be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, some of the most important method steps can be performed by such devices.

[0122] The inventive encoded audio signal may be stored on a digital storage medium or may be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet.

[0123] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. The implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, having stored thereon electronically readable control signals that cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Thus, the digital storage medium may be computer-readable.

[0124] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0125] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine-readable carrier.

[0126] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0127] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0128] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) on which is recorded the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0129] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.The data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, for example via the Internet.

[0130] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0131] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0132] A further embodiment according to the invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a storage device, etc. For example, the apparatus or system may comprise a file server for transmitting the computer program to the receiver.

[0133] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0134] The above embodiments are intended to illustrate the principles of the present invention only. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended that the present invention be limited only by the scope of the upcoming patent claims and not by the specific details presented through the description and explanation of the embodiments herein.

[0135] References

[0136]

[01] J.Le Roux and E.Vincente, "Consistent Wiener filtering for audiosource separation," IEEE Signal Processing Letters, pp.217-220, March 2013.

[0137]

[02] B.Jacob, J.Chen and EAPHabets, Speech enhancement in the STFTdomain, Springer Science & Business Media., 2011.

[0138]

[03] T. Virtanen, "Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria," IEEETRANS.ON AUDIO, SPEECH, AND LANGUAGE PROCES., pp.1066-1074, February 2007.

[0139]

[04] F.Weninger, JLRoux, JRHershey and S.Watanabe, "DiscriminativeNMF and its application to single-channel source separation," In FifteenthAnnual Conf.of the Intl.Speech Commun.Assoc., September 2014.

[0140]

[05] D.Wang and J.Chen,"Supervised speech separation based on deeplearning:An overview,"Proc.IEEE Intl.Conf.on Acoustics,Speech and SignalProcessing(ICASSP),pp.1702-1726,May 2018.

[0141]

[06] J.R.Hershey,Z.Chen,J.L.Roux and S.Watanabe,"Deep clustering:Discriminative embeddings for segmentation and separation,"Proc.IEEEIntl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),pp.31-35,March2016.

[0142]

[07] Y.Dong,M.Kolbaek,Z.H.Tan and J.Jensen,"Permutation invarianttraining of deep models for speaker-independent multi-talker speechseparation,"Proc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),pp.241-245,March 2017.

[0143]

[08] D.S.Williamson and D.Wang,"Speech dereverberation and denoisingusing complex ratio masks,"Proc.IEEE Intl.Conf.on Acoustics,Speech and SignalProcessing(ICASSP),pp.5590-5594,March 2017.

[0144]

[09] J.Lecomte et al.,"Packet-loss concealment technology advances inEVS,"Proc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),pp.5708-5712,August 2015.

[0145] [1]K.Han,Y.Wang,D.Wang,W.S.Woods,I.Merks,and T.Zhang,“Learningspectral mapping for speech dereverberation and denoising,”IEEE / ACMTrans.Audio,Speech,Lang.Process.,vol.23,no.6,pp.982–992,June 2015.

[0146] [2]Y.Wang,A.Narayanan,and D.Wang,“On training targets for supervisedspeech separation,”IEEE / ACM Trans.Audio,Speech,Lang.Process.,vol.22,no.12,pp.1849–1858,December 2014.

[0147] [3]D.S.Williamson,Y.Wang,and D.Wang,“Complex ratio masking formonaural speech separation,”IEEE Trans.Audio,Speech,Lang.Process.,vol.24,no.3,pp.483–492,March 2016.

[0148] [4]D.S.Williamson and D.Wang,“Speech dereverberation and denoisingusing complex ratio masks,”in Proc.IEEE Intl.Conf.on Acoustics,Speech andSignal Processing(ICASSP),March 2017,pp.5590–5594.

[0149] [5]J.R.Hershey,Z.Chen,J.L.Roux,and S.Watanabe,“Deep clustering:Discriminative embeddings for segmentation and separation,”in Proc.IEEEIntl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),March 2016,pp.31–35.

[0150] [6]Z.Chen,Y.Luo,and N.Mesgarani,“Deep attractor network for single-microphone speaker separation,”in Proc.IEEE Intl.Conf.on Acoustics,Speech andSignal Processing(ICASSP),March 2017,pp.246–250.

[0151] [7]Y.Isik,J.L.Roux,Z.Chen,S.Watanabe,and J.R.Hershey,“Single-channelmulti-speaker separation using deep clustering,”in Proc.Inter-speech Conf.,September 2016,pp.545–549.

[0152] [8]D.Yu,M. Z.H.Tan,and J.Jensen,“Permutation invariant trainingof deep models for speaker-independent multi-talker speech separation,”inProc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),March2017,pp.241–245.

[0153] [9]Y.Luo,Z.Chen,J.R.Hershey,J.L.Roux,and N.Mesgarani,“Deep clusteringand conventional networks for music separation:Stronger together,”inProc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),March2017,pp.61–65.

[0154]

[10] M.Kolbaek,D.Yu,Z.-H.Tan,J.Jensen,M.Kolbaek,D.Yu,Z.-H.Tan,andJ.Jensen,“Multitalker speech separation with utterance-level permutationinvariant training of deep recurrent neural networks,”IEEE Trans.Audio,Speech,Lang.Process.,vol.25,no.10,pp.1901–1913,October 2017.

[0155]

[11] W.Mack,S.Chakrabarty,F.-R. S.Braun,B.Edler,and E.A.P.Habets,“Single-channel dereverberation using direct MMSE optimization andbidirectional LSTM networks,”in Proc.Interspeech Conf.,September 2018,pp.1314–1318.

[0156]

[12] H.Erdogan and T.Yoshioka,“Investigations on data augmentation andloss functions for deep learning based speech-background separation,”inProc.Interspeech Conf.,September 2018,pp.3499–3503.

[0157]

[13] D.Wang,“On ideal binary mask as the computational goal of audi-tory scene analysis,”in Speech Separation by Humans and Machines,P.Divenyi,Ed.Kluwer Academic,2005,pp.181–197.

[0158]

[14] C.Hummersone,T.Stokes,and T.Brookes,“On the ideal ratio mask asthe goal of computational auditory scene analysis,”in Blind SourceSeparation,G.R.Naik and W.Wang,Eds.Springer,2014,pp.349–368.

[0159] [0]F.Mayer,D.S.Williamson,P.Mowlaee,and D.Wang,“Impact of phaseestimation on single-channel speech separation based on time-frequencymasking,”J.Acoust.Soc.Am.,vol.141,no.6,pp.4668–4679,2017.

[0160] [1]F.Weninger,H.Erdogan,S.Watanabe,E.Vincent,J.Roux,J.R.Hershey,andB.Schuller,“Speech enhancement with LSTM recurrent neural networks and itsapplication to noise-robust ASR,”in Proc.of the 12th Int.Conf.onLat.Var.An.and Sig.Sep.,ser.LVA / ICA.New York,USA:Springer-Verlag,2015,pp.91–99.

[0161] [2]X.Li,J.Li,and Y.Yan,“Ideal ratio mask estimation using deep neuralnetworks for monaural speech segregation in noisy reverberant conditions,”August 2017,pp.1203–1207.

[0162] [3]J.Benesty,J.Chen,and E.A.P.Habets,Speech Enhancement in the STFTDomain,ser.SpringerBriefs in Electrical and Computer Engineering.Springer-Verlag,2011.

[0163] [4]J.Benesty and Y.Huang,“A single-channel noise reduction MVDRfilter,”in Proc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),2011,pp.273–276.

[0164] [5]D.Fischer,S.Doclo,E.A.P.Habets,and T.Gerkmann,“Com-bined single-microphone Wiener and MVDR filtering based on speech interframe correlationsand speech presence probability,”in Speech Communication;12.ITG Symposium,Oct2016,pp.1–5.

[0165] [6]D.Fischer and S.Doclo,“Robust constrained MFMVDR filtering forsingle-microphone speech enhancement,”in Proc.Intl.Workshop Acoust.SignalEnhancement(IWAENC),2018,pp.41–45.

[0166] [7]S.Hochreiter and J.Schmidhuber,“Long short-term memory,”NeuralComputation,vol.9,no.8,pp.1735–1780,Nov 1997.

[0167] [8]J.B.D.Kingma,“Adam:A method for stochastic optimization,”inProc.IEEE Intl.Conf.on Learn.Repr.(ICLR),May 2015,pp.1–15.

[0168] [9]N.Srivastava,G.Hinton,A.Krizhevsky,I.Sutskever,andR.Salakhutdinov,“Dropout:A simple way to prevent neural networks fromoverfitting,”J.Mach.Learn.Res.,vol.15,no.1,pp.1929–1958,January 2014.[Online].Available: http: / / dl.acm.org / citation.cfm?id=2627435.2670313

[0169]

[10] J.F.Gemmeke,D.P.W.Ellis,D.Freedman,A.Jansen,W.Lawrence,R.C.Moore,M.Plakal,and M.Ritter,“Audio Set:An ontology and human-labeled dataset foraudio events,”in Proc.IEEE Intl.Conf.on Acoustics,Speech and SignalProcessing(ICASSP),March 2017,pp.776–780.

[0170]

[11] V.Panayotov,G.Chen,D.Povey,and S.Khudanpur,“Librispeech:An ASRcorpus based on public domain audio books,”in Proc.IEEE Intl.Conf.onAcoustics,Speech and Signal Processing(ICASSP),April 2015,pp.5206–5210.

[0171]

[12] C.Raffel,B.McFee,EJHumphrey,J.Salamon,O.Nieto,D.Liang,andD.PWEllis,“MIR EVAL:A transparent implementation of common MIR metrics,”inIntl.Soc.of Music Inf.Retrieval,October 2014,pp.367–372.

[0172]

[13] CHTaal,RCHendriks,R.Heusdens,and J.Jensen,“An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech,”IEEETrans.Audio,Speech,Long.Process.,vol.19,no.7,pp.2125–2136,September2011.

[0173]

[14] M.Relative,“pystoi,”https: / / github.com / relative / pystoi,2018.

Claims

1. A method for determining a depth filter (10x) for filtering a mixture of a desired signal and an undesired signal including an audio signal or a sensor signal to extract the desired signal from the mixture of the desired signal and the undesired signal, the method comprising the following steps: Determining (100) said depth filter (10x) having at least one dimension comprises: receiving (110) the mixture (10); estimating (120) the deep filter (10x) using a deep neural network, wherein the estimating (120) is performed using the mixture (10) and the desired representation (11) such that the deep filter (10x) when applied to an element of the mixture (10) obtains an estimate of an element of the desired representation (11), wherein the deep filter (10x) is obtained by defining a filter structure having filter variables for the deep filter (10x) having at least one dimension and training the deep neural network, wherein the training is performed using a mean square error (MSE) between a true value and the desired representation and minimizing the mean square error or minimizing an error function between the true value and the desired representation; wherein the depth filter (10x) has at least one dimension, the at least one dimension comprising elements (s x,y ) one- or multi-dimensional tensor.

2. The method according to claim 1, wherein the mixture (10) comprises a real-valued or complex-valued time-frequency representation or a feature representation thereof; and The expected representation (11) includes an expected real-valued or complex-valued time-frequency representation or a feature representation thereof.

3. The method according to claim 1, wherein the depth filter (10x) comprises a real-valued or complex-valued time-frequency filter; and / or wherein the depth filter (10x) having at least one dimension is described in a short-time Fourier transform domain.

4. The method according to claim 1, wherein the step of estimating (120) is performed for each element of the mixture (10) or for a predetermined portion of the elements of the mixture (10). The method of claim 1 , wherein the estimating ( 120 ) is performed for at least two sources. The method according to claim 1 , wherein the depth filter ( 10x ) is a multi-dimensional complex depth filter.

7. The method according to claim 1, wherein the deep neural network comprises a number of output parameters equal to the number of filter values ​​of the filter function of the deep filter (10x).

8. The method of claim 1, wherein the at least one dimension is selected from the group consisting of time, frequency, and sensor, or Wherein said at least one of said dimensions is across time or across frequency.

9. The method of claim 1, wherein the deep neural network comprises a batch normalization layer, a bidirectional long short-term memory layer, a feed-forward output layer with hyperbolic tangent activation, and / or one or more additional layers.

10. The method according to claim 1, further comprising the step of training the deep neural network.

11. The method according to claim 10, wherein the deep neural network is trained by optimizing the mean squared error between the true value of the desired representation (11) and an estimate of the desired representation (11); or wherein the deep neural network is trained by reducing a reconstruction error between the desired representation (11) and an estimate of the desired representation (11); or The training is performed by amplitude reconstruction.

12. The method of claim 1, wherein the estimating (120) is performed by using the following formula: where 2·L+1 is the filter dimension in the time frame direction, 2·I+1 is the filter dimension in the frequency direction, and is a complex conjugate one-dimensional or two-dimensional filter; and where is the expected representation of the estimate (11), where n is the time frame, k is the frequency index, and X(n,k) is the mixture.

13. The method of claim 10, wherein the training is performed by using the following formula: where X d (n,k) is the desired representation (11), and is the expected representation of the estimate (11), where N is the total number of time frames and K is the number of frequency bins per time frame, where n is the time frame and k is the frequency index, or By using the following formula: where X d (n,k) is the desired representation (11), and is the expected representation of the estimate (11), where N is the total number of time frames and K is the number of frequency bins per time frame, where n is the time frame and k is the frequency index.

14. The method according to claim 1, wherein the tensor elements (s x,y ) is bounded in magnitude or is bounded in magnitude by using the following formula: in is a complex conjugate two-dimensional filter.

15. The method of claim 1, wherein the step of applying is performed element-by-element.

16. The method according to claim 1, wherein the corresponding tensor elements (s x,y ) in the expectation representation (11) to execute the application.

17. The method according to claim 1, comprising a method (100) for filtering a mixture of a desired signal and an undesired signal comprising an audio signal or a sensor signal to extract the desired signal from the mixture of the desired signal and the undesired signal, the method comprising: The depth filter (10x) was applied to the mixture (10).

18. Application of the method (100) according to claim 17 for signal extraction or signal separation of at least two sources.

19. Use of the method (100) according to claim 17 for signal reconstruction.

20. Computer program product for performing one of the methods according to claim 1 when run on a computer.

21. An apparatus for determining a depth filter (10x) capable of extracting a desired signal from a mixture of a desired signal and an undesired signal, the apparatus comprising an input for receiving (110) said mixture (10) of said desired signal and said undesired signal, including an audio signal or a sensor signal, or a mixture (10) including at least the undesired signal; a deep neural network for estimating (120) the deep filter (10x) such that the deep filter (10x) when applied to the elements of the mixture (10) obtains an estimate of the elements of the desired representation (11); wherein, The deep filter (10x) is obtained by defining a filter structure having filter variables for a deep filter (10x) having at least one dimension and training the deep neural network, wherein the training is performed using a mean square error (MSE) between a true value and the desired representation and minimizing the mean square error or minimizing an error function between the true value and the desired representation; wherein the depth filter (10x) has at least one dimension, the at least one dimension comprising elements (s x,y ) one- or multi-dimensional tensor.

22. An apparatus for filtering a mixture, the apparatus comprising the device according to claim 21, the determined depth filter and means for applying the depth filter to the mixture.

Citation Information

Patent Citations

  • Channel environment adaptive OFDM receiving method based on neural network

    CN109194595A

  • Deep neural net based filter prediction for audio event classification and extraction

    US20160284346A1