Signal processing device, signal processing method, and program

By approximating the frequency covariance matrix with a linear combination of power spectrum and rank-1 covariance matrices, the signal processing device improves sound source separation for non-stationary signals, addressing the limitations of existing methods.

JP7709139B2Active Publication Date: 2025-07-16NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021203902
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-07-16
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

Existing sound source separation methods, such as IDLMA and PSDTF, fail to accurately model the inter-frequency correlation of source signals, particularly for non-stationary signals like voice and music, leading to suboptimal separation performance.

Method used

A signal processing device that approximates the frequency covariance matrix of source signals using a linear combination of a diagonal matrix representing power spectrum and a rank-1 frequency covariance matrix, estimated using deep neural networks, to improve separation performance.

Benefits of technology

The proposed method effectively models the frequency correlation of source signals, enhancing separation performance by accurately estimating the separation matrix, especially for non-stationary signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709139000014
    Figure 0007709139000014
  • Figure 0007709139000015
    Figure 0007709139000015
  • Figure 0007709139000016
    Figure 0007709139000016
Patent Text Reader

Abstract

To improve separation performance by properly approximating frequency covariance of a source signal.SOLUTION: A frequency covariance matrix approximating frequency correlation of a source signal is obtained and output by linear combination of a first matrix, which is a diagonal matrix representing an estimated power of the source signal in each frequency interval based on an observation signal based on a source signal emitted from one or more signal sources or a separation signal obtained by applying a separation filter to the observation signal, and a second matrix representing rank 1 frequency covariance corresponding to an estimated waveform of the source signal based on the observation signal or the separation signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to signal source separation technology. [Background technology]

[0002] A sound source separation technique that estimates the source signal of each sound source using an observed mixed acoustic signal as an input is a technique that is widely used for preprocessing of speech recognition, etc. Independent deep learned matrix analysis (IDLMA) is known as a method for performing sound source separation using multiple microphones (see, for example, Non-Patent Document 1, etc.).

[0003] Recently, independent positive semidefinite tensor analysis (IPSDTA) has been proposed as a blind source separation method that introduces positive semidefinite tensor factorization (PSDTF) into the sound source model (see, for example, Non-Patent Document 2, etc.). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Naoki Makishima et al., “Independent deeply learned matrix analysis for determined audio source separation.”, IEEE / ACM Transactions on Audio, Speech, and Language Processing 27.10 (2019): 1601-1615, [Retrieved February 5, 2021], Internet<https: / / ieeexplore.ieee.org / stamp / stamp.jsp?tp=&arnumber=<8747523> [Non-Patent Document 2] R. Ikeshita, “Independent positive semidefinite tensor analysis in blind source separation,” in Proc. EUSIPCO, pp. 1652-1656, Sep. 2018.

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the probabilistic model of IDLMA does not consider the inter-frequency correlation of source signals. Therefore, the separation performance of IDLMA for non-stationary signals such as voice signals and music signals is low.

[0006] Also, in the model of PSDTF, the frequency covariance matrix is expressed as a convex cone combination of time-invariant bases (positive semi-definite matrices). Although this combination weight is time-varying, since the frequency correlation changes according to the time frame, the model of PSDTF that expresses the frequency covariance by the convex cone combination of time-invariant bases is inaccurate. Therefore, there is still room for improvement in the separation performance of PSDTF.

[0007] Such problems are not limited to the source separation of acoustic signals, but are common to source separation that takes an observation signal based on source signals emitted from multiple signal sources as input and estimates the source signals of each signal source.

[0008] The present invention has been made in view of such points, and an object thereof is to appropriately model the frequency covariance of source signals and improve the separation performance.

Means for Solving the Problems

[0009] The signal processing device obtains a frequency covariance matrix that approximates the frequency correlation of the source signal by a linear combination of a first matrix, which is a diagonal matrix representing the estimated power of the source signal in each frequency band based on an observation signal based on a source signal emitted from one or more signal sources or a separated signal obtained by applying a separation filter to the observation signal, and a second matrix, which represents a rank-1 frequency covariance corresponding to the estimated waveform of the source signal based on the observation signal or the separated signal.

Advantages of the Invention

[0010] Thereby, the frequency covariance of the source signal can be appropriately approximated to improve the separation performance.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described. [Principle] First, the principle of this embodiment will be described. Assume an environment where time-series source signals emitted from N signal sources propagate in space and are observed by M sensors, thereby obtaining observed signals in the time domain. Here, N and M are positive integers satisfying N ≤ M. Practically meaningful source separation is the case where N ≥ 2. In this case, a mixed signal in which the source signals are mixed in space is observed by the sensors. The time-series signal observed by the sensors is called the observed signal in the time domain. Usually, the observed signal in the time domain is a digital signal converted by analog-to-digital conversion. Also, the source signal estimated from the observed signal is called the separated signal. Examples of the source signal, observed signal, and separated signal include acoustic signals (such as voice signals and music signals), ultrasonic signals, electromagnetic wave signals, biological signals, seismic wave signals, etc. For example, when the source signal, observed signal, and separated signal are acoustic signals, the signal source is a sound source, the source signal is an acoustic signal, and the sensor is a microphone.

[0013] The source signal, observed signal, and separated signal in the time domain are converted into the source signal s ij , observed signal x ij , and separated signal y ij in the time-frequency domain obtained by converting each predetermined time window by a short-time Fourier transform (STFT) or the like, and are represented as follows. s ij =(s ij1 ,..., s ijn ,..., s ijN ) T ∈C N (1) x ij =(x ij1 ,..., x ijm ,..., x ijM ) T ∈C M (2) y ij =(y ij1 ,..., y ijn ,..., y ijN ) T ∈C N (3) Here, i = 1, ..., I, j = 1, ..., J, n = 1, ..., N, and m = 1, ..., M represent indexes that identify a predetermined frequency interval (e.g., frequency bin, discrete frequency), a predetermined time interval (e.g., time frame), a signal source (e.g., sound source), and an observation channel (e.g., microphone), respectively. I, J, N, and M represent positive integers, and β T represents the transposition of β (e.g., matrix or vector), and C represents the entire set of complex numbers. That is, the source signal s ij , the observation signal x ij , and the separated signal y ij are time-frequency domain signals, and β jn represents the element corresponding to the time interval and signal source represented by j and n, and β ijn represents the element corresponding to the frequency interval, time interval, and signal source represented by i, j, and n.

[0014] Assuming that the mixing system in space is time-invariant and the reverberation time is shorter than the time window length, an approximation of x i ∈C M×N using the time-invariant mixing matrix A ij = A i s ij holds. Assuming that A i is regular, the separation matrix (separation filter) W i = (w i1 , ..., w in , ..., w iN ) H = A i -1 exists, and the separated signal is y ij = W i x ij . Even if A i is not regular, the Moore-Penrose pseudo-inverse matrix (generalized inverse matrix) of A i can be used as the separation matrix W i . Here, β H represents the Hermitian transposition of β. At this time, assuming the independence of the separated signals, the negative log-likelihood of the observation signals can be expressed as follows.

Equation

Number

Number

[0015] The probability distribution model p(Y n ) is expressed as follows.

Number

[0016] D jn is the observed signal x ij =(x ij1 ,...,x ijm ,...,x ijM ) T Also, the separated signal y ij =(y ij1,...,y ijn ,...,y ijN ) T The source signal s in each frequency interval (e.g., each frequency bin) based on ij =(s ij1 ,...,s ijn ,...,s ijN ) T is a diagonal matrix (first matrix) representing the estimated power (e.g., power spectrum) of jn In this example, D jn is an I×I diagonal matrix. For example, D 1jn 2 ,…,d Ijn 2 The I×I diagonal matrix D with diagonal components jn =diag{d 1jn 2 ,…,d Ijn 2}. Here, d ijn is a non - negative real number. d ijn 2 is the estimated power (e.g., power spectrum) of the source signal s in each frequency interval (e.g., each frequency bin) based on the observed signal x ij or the separated signal y ij It is a scale parameter. For example, the estimated power of the source signal s in each frequency interval ij or D ij is estimated using a DNN (Deep Neural Network) (first deep neural network) based on the observed signal x jn or the separated signal y ij or the separated signal y ij By this, the estimation accuracy of D jn is improved, thereby obtaining a properly approximated frequency covariance matrix R jn and as a result, the separation performance can be improved. For example, D ij can be estimated by inputting the time - domain signal of the observed signal x ij or the separated signal y ij or the separated signal y ij into a pre - trained DNN. For example, D jn can be estimated by inputting |Y n | into a pre - trained DNN.jn can be obtained (see, for example, Non-Patent Document 1, etc.). Here, |Y n | represents an I×J matrix with the absolute value of each element y n of Y ijn (i.e., |y ijn |) as the (i,j) element. However, this does not limit the present invention. The separation signal y ij = (y ij1 ,..., y ijN ) T can be used as it is to set D jn = diag{d 1jn 2 , …, d Ijn 2} = diag{y 1jn 2 , …, y Ijn 2} as shown. jn That is, D can be set in this way.

[0017] (v jn → )(v jn → ) H represents the rank-1 frequency covariance corresponding to the estimated waveform of the source signal s ij = (x ij1 ,..., x ijm ,..., x ijM ) T or the separation signal y ij = (y ij1 ,..., y ijn ,..., y ijN ) T (the second matrix). In this example, (v ij = (s ij1 ,..., s ijn ,..., s ijN ) T (v jn → )(v jn → ) H is an I×I matrix. For example, v jn → is an I-dimensional vector v 1jn , …, v Ijn with elements v jn→ =(v 1jn ,…,v Ijn ) T ∈C I is. For example, the estimated waveform v ij of the source signal s jn → is estimated using a DNN (second deep neural network) based on the observation signal x ij or the separated signal y ij . Thereby, the estimation accuracy of (v jn → )(v jn → ) H can be improved, and thereby an appropriately approximated frequency covariance matrix R jn can be obtained, and as a result, the separation performance can be improved. For example, the observation signal x ij or the separated signal y ij or the time-domain signal of the observation signal x ij or the separated signal y ij is input to a pre-trained DNN to estimate v jn → . However, this does not limit the present invention, and (v ij or the separated signal y ij or the time-domain signal of the observation signal x ij or the separated signal y ij is input to a pre-trained DNN, (v jn → )(v jn → ) H may be directly obtained. Alternatively, the separated signal y ij =(y ij1 ,...,y ijN ) T is used as it is to set v jn → =(v 1jn ,…,v Ijn )=(y 1jn ,…,y Ijn ) like v jn → .

[0018] α1 (the first weight) and α2 (the second weight) are weights (coefficients), and two matrices D jn and (v jn → )(v jn → ) H represent the mixing ratio. That is, the frequency covariance matrix R jn represented by Equation (6) is the diagonal matrix D jn (the first matrix) weighted by α1 (the first weight), α1D jn (the first weighted matrix), and the matrix (v jn → )(v jn → ) H (the second matrix) weighted by α2 (the second weight), α2(v jn → )(v jn → ) H (the second weighted matrix), and is a linear combination (e.g., sum of products). α1 and α2 are real numbers. For example, when α is a non - negative real number, α1 = 1 - α and α2 = α, the frequency covariance matrix R jn is a convex combination of two matrices D jn and (v jn → )(v jn → ) H By setting the coefficient α2 to be non - zero (e.g., positive), the frequency correlation of the source signal s jn is reflected in R ij , and it becomes possible to estimate the separation matrix W ij considering the frequency correlation of the source signal s i . On the other hand, when the coefficient α1 is 0 and the coefficient α2 is not 0, only (v jn jn → )(v jn → H ) jn is reflected in R jn , and problems may occur such that R jn loses rank and an appropriate solution for the separation matrix cannot be obtained, or the cost function for separation matrix estimation becomes unstable. Here, by setting the coefficient α1 to be non - zero (e.g., positive), D jnis reflected. D which is a diagonal matrix jn will not have a rank drop, and R jn when D jn is reflected, the cost function for separating matrix estimation is less likely to become unstable. On the other hand, when coefficient α1 is not zero and coefficient α2 is zero, and R jn when only D jn is reflected, the frequency correlation of the source signal s jn in R ij is not reflected, and the separation matrix estimation considering the frequency correlation of the source signal s ij cannot be performed. Therefore, it is desirable that both α1 and α2 are non - zero (for example, positive). For example, when α1 = 1 - α and α2 = α, it is desirable that 0 < α < 1. By setting both α1 and α2 to be non - zero in this way, while suppressing the rank drop of R jn the frequency correlation of the source signal s ij is considered, and a separation matrix W i with high separation performance can be estimated.

[0019] In this embodiment, the frequency covariance matrix R jn exemplified by Equation (6), that is, the diagonal matrix (the first matrix) D ij representing the estimated power of the source signal s ij in each frequency interval based on the observed signal x ij or the separated signal y jn and the matrix (the second matrix) (v ij representing the rank - 1 frequency covariance corresponding to the estimated waveform of the source signal s ij based on the observed signal x ij or the separated signal y jn → )(v jn → ) H and, the frequency covariance matrix R jn that approximates the frequency correlation of the source signal by a linear combination of and, the observed signal x ij and, based on which the separation matrix W i is updated by an optimization process. Although not limited to this optimization process, for example, as the frequency covariance matrix, the above - mentioned frequency covariance matrix R jnIt may be performed according to the same procedure as IPSDTA (see, for example, Non-Patent Document 2, etc.). For example, for the frequency covariance matrix R jn and the observation signal x ij by applying the processing of the following formulas (7)-(14), the separation matrix W i =(w i1 ,...,w in ,...,w iN ) H is updated.

Equation

Equation

Equation

Number

Number

Number

Number

Number

Number

[0020] To summarize, in this embodiment, the estimated power of the source signal s ij or the separated signal y ij in each frequency interval based on the ij The diagonal matrix (first matrix) D representing jn and the observed signal x ij or the separated signal yij The source signal s based on ij The matrix (second matrix) (v representing the rank-1 frequency covariance corresponding to the estimated waveform of jn → )(v jn → ) H and the frequency covariance matrix R that approximates the frequency correlation of the source signal by a linear combination of jn are used. The two matrices D jn constituting the frequency covariance matrix R jn and (v jn → )(v jn → ) H are both time-varying, and their linear combination R jn has a high degree of freedom in expression and can be appropriately approximated to the frequency covariance of the source signal s ij . By updating the separation matrix W jn based on such a frequency covariance matrix R i , an appropriate separation matrix W ij can be obtained regardless of whether the source signal s i is stationary or non-stationary, and the separation performance is improved.

[0021] Also, the diagonal matrix D jn (first matrix) weighted by a non-zero α1 (first weight) to obtain α1D jn (first weighted matrix), and the matrix (v representing the rank-1 frequency covariance jn → )(v jn → ) H (second matrix) weighted by a non-zero α2 (second weight) to obtain α2(v jn → )(v jn → ) H (second weighted matrix), and by expressing the frequency covariance matrix R jn as a linear combination of these, the frequency correlation of the source signal s ij can be considered, and the instability of the cost function can be suppressed, and a separation matrix W i with high separation performance can be obtained.

[0022] Also, the estimated power of the source signal s in each frequency band or D ij is estimated using a DNN (first deep neural network) based on the observed signal x jn or the separated signal y ij or v ij is estimated using a DNN (first deep neural network) based on the observed signal x jn → or (v jn → )(v jn → ) H is estimated using a DNN (second deep neural network) based on the observed signal x ij or the separated signal y ij to improve the separation performance.

[0023] [First Embodiment] Next, the first embodiment of the present invention will be described. [Configuration] As illustrated in FIG. 1, the signal processing apparatus 1 of the present embodiment includes a storage unit 101, an initial value setting unit 11, a signal separation unit 12, a frequency covariance matrix estimation unit 13, a separation filter estimation unit 14, and a control unit 15. Further, the signal processing apparatus 1 may include a time-frequency domain conversion unit 102. Here, the frequency covariance matrix estimation unit 13 illustrated includes a time domain conversion unit 131, a waveform estimation unit 132, a time-frequency domain conversion unit 133, a power estimation unit 134, and a combining unit 135. The signal processing apparatus 1 executes each process under the control of the control unit 15, and the data obtained in each process is sequentially stored in the storage unit 101 and read out as needed and used in each process.

[0024] [Pretreatment] As a prerequisite for the signal separation process, the observed signal x in the time-frequency domain ij =(x ij1 ,...,x ijm ,...,x ijM ) T to be separated is stored in the storage unit 101. For example, the observed signal x in the time-frequency domain ij is input to the signal processing apparatus 1 and stored in the storage unit 101. Alternatively, the observed signal X in the time domain pis input to the signal processing device 1, and the time-frequency domain conversion unit 102 converts the observed signal X in the time domain into an observed signal x in the time-frequency domain by, for example, short-time Fourier transform or the like and stores it in the storage unit 101. Here, p = 1,..., P is an index representing a predetermined time interval (for example, a time frame), P is an integer of 1 or more, and X p represents the observed signal in the time-frequency domain in the time interval represented by p. ij Note that p = 1,..., P is an index representing a predetermined time interval (for example, a time frame), P is an integer of 1 or more, and X p represents the observed signal in the time-frequency domain in the time interval represented by p.

[0025] <Signal separation process> The signal separation process of the present embodiment will be illustrated with reference to FIGS. 2 and 3.

[0026] <Processing of the initial setting unit 11 (step S11)> The initial value setting unit 11 sets values for the above-described weights α1 and α2. For example, when α1 = 1 - α and α2 = α, the initial value setting unit 11 sets a value for α. Information on the values of the weights α1 and α2 is sent to the integration unit 135 of the frequency covariance matrix estimation unit 13. Further, the initial value setting unit 11 sets an initial value for the separation matrix W i . An example of the initial value of the separation matrix W i is the identity matrix. Information on the initial value of the separation matrix W i is sent to the separation filter estimation unit 14.

[0027] <Processing of the signal separation unit 12 (step S12)> The observed signal x ij read from the storage unit 101 and the separation matrix W i output from the separation filter estimation unit 14 are input to the signal separation unit 12. The signal separation unit 12 applies the separation matrix W ij (separation filter) to the input observed signal x i (observed signal based on source signals emitted from one or more signal sources) to obtain and output a separated signal y ij . For example, the signal separation unit 12 uses the input observed signal x ij and the separation matrix W i to obtain and output a separated signal y ij according to the following formula (20). yij =W i x ij (20)

[0028] <Processing of Frequency Covariance Matrix Estimation Unit 13 (Step S13)> The frequency covariance estimation unit 13 receives the separated signal y obtained in step S12. ij The frequency covariance estimation unit 13 of the present embodiment is based on the separated signal y. ij The diagonal matrix (first matrix) D representing the estimated power of the source signal s in each frequency interval. ij And the matrix (second matrix) (v representing the rank-1 frequency covariance corresponding to the estimated waveform of the source signal s based on the separated signal y. jn Based on the separated signal y. ij The source signal s. ij The matrix (v representing the rank-1 frequency covariance corresponding to the estimated waveform of the source signal s. jn → )(v jn → ) H And the frequency covariance matrix R that approximates the frequency correlation of the source signal by a linear combination of and is obtained and output. The details of this process are illustrated using FIG. 3. The separated signal y is input to the time domain conversion unit 131. jn The time domain conversion unit 131 converts the separated signal y into the separated signal Γ in the time domain by, for example, inverse short-time Fourier transform (ISTFT) and outputs it (step S131). The separated signal Γ in the time domain is input to the waveform estimation unit 132. Here, a DNN (DNN for waveform estimation) for estimating the waveform of the source signal in the time domain from the source signal or separated signal in the time domain is pre-learned, and this waveform estimation DNN is set in the waveform estimation unit 132. The waveform estimation unit 132 inputs the separated signal Γ in the time domain to the waveform estimation DNN and infers (estimates) and outputs the estimated waveform V of the source signal in the time domain (step S132). The estimated waveform V in the time domain is input to the time-frequency domain conversion unit 133. The time-frequency domain conversion unit 133 converts the estimated waveform V in the time domain into, for example, by short-time Fourier transform. ij Is input. The time domain conversion unit 131 converts the separated signal y into the separated signal Γ in the time domain by, for example, inverse short-time Fourier transform (ISTFT) and outputs it (step S131). ij Into the separated signal Γ in the time domain and outputs it (step S131). The separated signal Γ in the time domain. p Is converted and output (step S131). The separated signal Γ in the time domain. p Is input to the waveform estimation unit 132. Here, a DNN (DNN for waveform estimation) for estimating the waveform of the source signal in the time domain from the source signal or separated signal in the time domain is pre-learned, and this waveform estimation DNN is set in the waveform estimation unit 132. p The separated signal Γ in the time domain is input to the waveform estimation DNN, and the estimated waveform V of the source signal in the time domain is inferred (estimated) and output (step S132). p Is inferred (estimated) and output (step S132). The estimated waveform V in the time domain. p Is input to the time-frequency domain conversion unit 133. The time-frequency domain conversion unit 133 converts the estimated waveform V in the time domain into, for example, by short-time Fourier transform.p is converted into an estimated waveform v in the time-frequency domain jn → and output (step S133). Further, the separation signal y ij is input to the power estimation unit 134. Here, a DNN (DNN for power estimation) for estimating the power (e.g., power spectrum) of the source signal in each frequency interval from the source signal or separation signal in the time-frequency domain has been pre-trained, and this DNN for power estimation is set in the power estimation unit 134. The power estimation unit 134 inputs the separation signal y ij to the DNN for power estimation, and infers (estimates) and outputs a matrix D ij representing the estimated power of the source signal s ij in the time-frequency domain (step S134). The estimated waveform v jn → obtained in step S133 ij and the matrix D ij obtained in step S134 are input to the combining unit 135. The combining unit 135 linearly combines the matrix D jn → and the matrix (v jn → )(v H using the aforementioned weights α1 and α2 as in equation (6) to obtain and output a frequency covariance matrix R jn (step S135).

[0029] <Processing of the separation filter estimation unit 14 (step S14)> The frequency covariance matrix R jn obtained in step S13 and the observation signal x ij read from the storage unit 101 are input to the separation filter estimation unit 14. The separation filter estimation unit 14 updates and outputs a separation matrix W jn based on the frequency covariance matrix R ij and the observation signal x i . For example, the separation filter estimation unit 14 optimizes and outputs the separation matrix W jn while fixing the frequency covariance matrix R i obtained in step S13. For example, the separation filter estimation unit 14 follows the aforementioned equations (7)-(14) to optimize the separation matrix W iUpdate and output it (Step S14).

[0030] <Processing of control unit 15 (Step S15)> The control unit 15 determines whether the update process of the separation matrix W i has satisfied the termination condition. There is no limitation on the termination condition. For example, the processes from Step S12 to S15 have been repeated a predetermined number of times or more, the frequency covariance matrix R jn or the separation matrix W i has an update amount less than or equal to a predetermined value, etc., can be used as the termination condition. If it is determined that the termination condition is not satisfied, the process returns to Step S12. On the other hand, if it is determined that the termination condition is satisfied, the process proceeds to the following Step S16.

[0031] <Output processing (Step S16)> When the termination condition is satisfied, the signal processing device 1 uses the separated signal Γ in the time domain obtained in Step S131 of Step S13 p , or the separated signal y in the time-frequency domain obtained in Step S12 ij to output.

[0032] [Modification Example 1 of the First Embodiment] In the first embodiment, a DNN for waveform estimation for estimating the waveform of the source signal in the time domain from the source signal or separated signal in the time domain is pre-trained, and the waveform estimation unit 132 inputs the separated signal Γ in the time domain p into this DNN for waveform estimation, and infers and outputs the estimated waveform V of the source signal in the time domain p . Instead of this, a DNN for waveform estimation for estimating the waveform of the source signal in the time domain from the source signal or separated signal in the time-frequency domain is pre-trained, and the waveform estimation unit 132 inputs the separated signal y in the time-frequency domain ij into the DNN for waveform estimation, and infers and outputs the estimated waveform V of the source signal in the time domain p may be.

[0033] Alternatively, a DNN for waveform estimation for estimating the waveform of the source signal in the time-frequency domain from the source signal or separated signal in the time-frequency domain is pre-trained, and the waveform estimation unit 132 inputs the separated signal y in the time-frequency domainij is input into the DNN for waveform estimation, and the estimated waveform v in the time-frequency domain jn → may be inferred and output. In this case, the time-frequency domain conversion unit 133 can be omitted, and the estimated waveform v output from the waveform estimation unit 132 jn → is input into the combining unit 135.

[0034] Alternatively, the separated signal y obtained in step S12 ij =(y ij1 ,...,y ijn ,...,y ijN ) T is used as the estimated waveform v jn → =(v 1jn ,…,v Ijn ) T =(y 1jn ,…,y Ijn ) T and input into the combining unit 135. In this case, the waveform estimation unit 132 and the time-frequency domain conversion unit 133 can be omitted.

[0035] Alternatively, a DNN for waveform estimation for estimating the matrix (v jn → )(v jn → ) H from the source signal or separated signal in the time domain or time-frequency domain is pre-trained, and the waveform estimation unit 132 inputs the separated signal Γ in the time domain p or the separated signal y in the time-frequency domain ij into the DNN for waveform estimation, and infers and outputs the matrix (v jn → )(v jn → ) H may be. In this case, the time-frequency domain conversion unit 133 can be omitted, and the matrix (v jn → )(v jn → ) H output from the waveform estimation unit 132 is input into the combining unit 135.

[0036] [Modification Example 2 of the First Embodiment] In the first embodiment, a DNN for power estimation for estimating the power of the source signal in each frequency band from the source signal or the separated signal in the time-frequency domain is pre-trained, and the power estimation unit 134 separates the signal y in the time-frequency domain ij is input to this waveform estimation DNN, and the matrix D ij is inferred and output. Instead of this, a DNN for power estimation for estimating the power of the source signal in each frequency band from the source signal or the separated signal in the time domain is pre-trained, and the power estimation unit 134 separates the signal Γ in the time domain p is input to this waveform estimation DNN, and the matrix D ij may be inferred and output.

[0037] [Second Embodiment] Next, a second embodiment of the present invention will be described. In the first embodiment, the frequency covariance matrix R ij was obtained based on the separated signal y jn However, the frequency covariance matrix R ij may be obtained based on the observation signal x jn Hereinafter, the description will focus on the differences from the first embodiment, and the same reference numerals will be used for the common matters to simplify the description.

[0038] [Configuration] As illustrated in FIG. 1, the signal processing apparatus 2 of the present embodiment includes a storage unit 101, an initial value setting unit 11, a signal separation unit 12, a frequency covariance matrix estimation unit 23, a separation filter estimation unit 14, and a control unit 15. Further, the signal processing apparatus 2 may include a time-frequency domain conversion unit 102. The frequency covariance matrix estimation unit 23 illustrated here includes a time domain conversion unit 131, a waveform estimation unit 232, a time-frequency domain conversion unit 133, a power estimation unit 234, and a coupling unit 135. The signal processing apparatus 2 executes each process under the control of the control unit 15, and the data obtained in each process is sequentially stored in the storage unit 101 and read out as necessary and used in each process.

[0039] [Preprocessing] It is the same as the first embodiment except that the frequency covariance matrix estimation unit 13 is replaced with the frequency covariance matrix estimation unit 23.

[0040] <Signal separation process> Using FIGS. 2 and 4, the signal separation process of this embodiment is illustrated. First, after the process of step S12 described in the first embodiment, the following process of the frequency covariance matrix estimation unit 23 (step S23) is executed.

[0041] <Process of frequency covariance matrix estimation unit 23 (step S23)> The frequency covariance estimation unit 23 is input with the observation signal x in the time-frequency domain read from the storage unit 101. ij The frequency covariance estimation unit 23 of this embodiment is based on the observation signal x. ij The diagonal matrix (first matrix) D representing the estimated power of the source signal s in each frequency interval. ij And the matrix (second matrix) (v jn Based on the observation signal x. ij The frequency covariance of rank 1 corresponding to the estimated waveform of the source signal s. ij (v jn → )(v jn → ) H And a frequency covariance matrix R that approximates the frequency correlation of the source signal by a linear combination of jn Is obtained and output. Using FIG. 4, the details of this process are illustrated. The observation signal x is input to the time domain conversion unit 131. ij The time domain conversion unit 131 converts the observation signal x ij Into the observation signal X in the time domain by, for example, inverse short-time Fourier transform or the like and outputs it (step S231). The observation signal X in the time domain p Is input to the waveform estimation unit 232. Here, a DNN (DNN for waveform estimation) for estimating the waveform of the source signal in the time domain from the observation signal in the time domain has been pre-learned, and this DNN for waveform estimation is set in the waveform estimation unit 232. The waveform estimation unit 232 inputs the observation signal X in the time domain p To the DNN for waveform estimation, and the estimated waveform V of the source signal in the time domain. p And outputs it.p Infer (estimate) and output it (step S232). The estimated waveform V in the time domain p is input to the time-frequency domain conversion unit 133. The time-frequency domain conversion unit 133 converts the estimated waveform V in the time domain into an estimated waveform v in the time-frequency domain by, for example, short-time Fourier transform, etc., and outputs it (step S133). Also, the observation signal x p is input to the power estimation unit 234. Here, a DNN (DNN for power estimation) for estimating the power (for example, power spectrum) of the source signal s in each frequency interval from the observation signal in the time-frequency domain has been pre-learned, and this DNN for power estimation is set in the power estimation unit 234. The power estimation unit 234 inputs the observation signal x jn → to the DNN for power estimation, and infers (estimates) and outputs a matrix D ij representing the estimated power of the source signal s in the time-frequency domain (step S234). The estimated waveform v ij obtained in step S133 and the matrix D ij obtained in step S234 are input to the combining unit 135. The combining unit 135 linearly combines the matrix D ij and the matrix (v ij )(v jn → ) ij by the aforementioned weights α1 and α2 as in Equation (6) to obtain and output a frequency covariance matrix R ij and (v jn → )(v jn → ) H to obtain and output a frequency covariance matrix R jn (step S135).

[0042] Other processes are the same as those in the first embodiment.

[0043] [Modification Example 1 of the Second Embodiment] In the second embodiment, a waveform estimation DNN for estimating the waveform of the source signal in the time domain from the observation signal in the time domain has been pre-learned, and the waveform estimation unit 232 inputs the observation signal X p in the time domain to this waveform estimation DNN, and the estimated waveform V pwas inferred and output. Instead of this, a DNN for waveform estimation for estimating the waveform of the source signal in the time domain from the observed signal in the time-frequency domain is pre-trained, and the waveform estimation unit 232 inputs the observed signal x in the time-frequency domain ij to the DNN for waveform estimation, and may infer and output the estimated waveform V of the source signal in the time domain p .

[0044] Alternatively, a DNN for waveform estimation for estimating the waveform of the source signal in the time-frequency domain from the observed signal in the time-frequency domain is pre-trained, and the waveform estimation unit 232 inputs the observed signal x in the time-frequency domain ij to the DNN for waveform estimation, and may infer and output the estimated waveform v in the time-frequency domain jn → . In this case, the time-frequency domain conversion unit 133 can be omitted, and the estimated waveform v output from the waveform estimation unit 232 jn → is input to the combining unit 135

[0045] Alternatively, from the observed signal in the time domain or the time-frequency domain, a DNN for waveform estimation for estimating the matrix (v jn → )(v jn → ) H is pre-trained, and the waveform estimation unit 232 inputs the observed signal X in the time domain p or the observed signal x in the time-frequency domain ij to the DNN for waveform estimation, and may infer and output the matrix (v jn → )(v jn → ) H . In this case, the time-frequency domain conversion unit 133 can be omitted, and the matrix (v jn → )(v jn → ) H output from the waveform estimation unit 232 is input to the combining unit 135

[0046] [Modification Example 2 of the Second Embodiment] In the second embodiment, a DNN for power estimation for estimating the power of the source signal in each frequency interval from the observation signal in the time-frequency domain is pre-trained, and the power estimation unit 234 separates the signal y in the time-frequency domain ij is input to this waveform estimation DNN, and the matrix D ij is inferred and output. Instead of this, a DNN for power estimation for estimating the power of the source signal in each frequency interval from the observation signal in the time domain is pre-trained, and the power estimation unit 234 inputs the observation signal X in the time domain p to the waveform estimation DNN, and the matrix D ij may be inferred and output.

[0047] [Evaluation experiment results] The signal source separation (sound source separation) experiment results when the source signal is a music signal are shown. In the experiment, three methods were compared: the method (FSCM+DNN) described in Reference 1, the method (IDLMA) described in Non-Patent Document 1, and the method (IDLTA) of the first embodiment. In each method, the separation matrix W i was initialized with the identity matrix, and the number of iterations of the separation matrix update process was set to 100. In the experiment, two-source separation (N=2) was performed, and three combinations of mixed sounds of bass (Ba.) and vocal (Vo.), Vo. and drum (Dr.), and Dr. and Ba. were used as observation signals for evaluation. In addition, the amount of improvement in the source-to-distortion ratio (SDR) was used as an index of separation performance.

[0048] FIG. 5A shows the average value of the separation performance when α is changed in the range of 0≦α≦1 for the IDLTA, which is the method described in Non-Patent Document 1, with α1=1-α and α2=α, the diagonal matrix D jn and the matrix (v jn → )(v jn → ) H representing the frequency covariance of rank 1. The horizontal axis in FIG. 5A represents the value of α, and the vertical axis represents the SDR improvement amount [dB]. In this example, it can be seen that as the value of α increases from 0, the separation performance improves, and the highest performance is shown around α=0.5. Also, it can be seen that the separation performance decreases when α approaches 1.

[0049] Figure 5B shows the separation performance of each comparison method. The horizontal axis of Figure 5B represents the type of observed signal (mixed sound of Ba. and Vo., mixed sound of Vo. and Dr., mixed sound of Dr. and Ba.), and the vertical axis represents the amount of SDR improvement [dB]. Here, α = 0.5. In the case of any observed signal, IDLTA exceeds the separation performance of other comparison methods, and the superiority of the method of this embodiment can be confirmed.

[0050] [Hardware Configuration] The signal processing devices 1 and 2 in each embodiment are devices configured by a general-purpose or dedicated computer including a processor (hardware processor) such as a CPU (central processing unit) and a memory such as a RAM (random-access memory) and a ROM (read-only memory) executing a predetermined program. That is, the signal processing devices 1 and 2 in each embodiment have, for example, a processing circuitry configured to implement each part they have. This computer may include one processor and memory, or may include a plurality of processors and memory. This program may be installed in the computer or may be recorded in a ROM or the like in advance. Also, instead of an electronic circuit that realizes a functional configuration by loading a program like a CPU, a part or all of the processing units may be configured using an electronic circuit that realizes a processing function alone. Also, the electronic circuit constituting one device may include a plurality of CPUs.

[0051] FIG. 6 is a block diagram illustrating the hardware configurations of the signal processing devices 1 and 2 in each embodiment. As illustrated in FIG. 6, the signal processing devices 1 and 2 in this example have a CPU (Central Processing Unit) 10a, an input unit 10b, an output unit 10c, a RAM (Random Access Memory) 10d, a ROM (Read Only Memory) 10e, an auxiliary storage device 10f, and a bus 10g. The CPU 10a in this example has a control unit 10aa, an arithmetic unit 10ab, and a register 10ac, and executes various arithmetic processes according to various programs read into the register 10ac. The input unit 10b is an input terminal, a keyboard, a mouse, a touch panel, etc. to which data is input. The output unit 10c is an output terminal to which data is output, a display, a LAN card, etc. controlled by the CPU 10a that has read a predetermined program. The RAM 10d is an SRAM (Static Random Access Memory), a DRAM (Dynamic Random Access Memory), etc., and has a program area 10da in which a predetermined program is stored and a data area 10db in which various data is stored. The auxiliary storage device 10f is, for example, a hard disk, an MO (Magneto-Optical disc), a semiconductor memory, etc., and has a program area 10fa in which a predetermined program is stored and a data area 10fb in which various data is stored. The bus 10g connects the CPU 10a, the input unit 10b, the output unit 10c, the RAM 10d, the ROM 10e, and the auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f into the program area 10da of the RAM 10d according to the read OS (Operating System) program. Similarly, the CPU 10a writes the various data stored in the data area 10fb of the auxiliary storage device 10f into the data area 10db of the RAM 10d. Then, the address on the RAM 10d into which this program and data are written is stored in the register 10ac of the CPU 10a.The control unit 10aa of the CPU 10a sequentially reads out these addresses stored in the register 10ac, reads out programs and data from the area on the RAM 10d indicated by the read-out addresses, sequentially causes the arithmetic unit 10ab to execute the operations indicated by the programs, and stores the operation results in the register 10ac. With such a configuration, the functional configurations of the signal processing devices 1 and 2 are realized.

[0052] The above-described program can be recorded on a computer-readable recording medium. Examples of computer-readable recording media are non-transitory recording media. Examples of such recording media are magnetic recording devices, optical disks, magneto-optical recording media, semiconductor memories, and the like.

[0053] The distribution of this program can be carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Further, the program may be stored in the storage device of a server computer, and configured to be distributed by transferring the program from the server computer to other computers via a network. As described above, a computer that executes such a program first stores, for example, the program recorded on a portable recording medium or the program transferred from a server computer in its own storage device once. Then, at the time of executing processing, this computer reads the program stored in its own storage device and executes processing according to the read program. Also, as another execution form of this program, it is also possible that the computer directly reads the program from the portable recording medium and executes processing according to the program. Further, each time a program is transferred to this computer from a server computer, it is also possible to sequentially execute processing according to the received program. Also, it is also possible to configure to execute the above-described processing by a so-called ASP (Application Service Provider) type service that realizes a processing function only by an execution instruction and result acquisition without transferring the program from the server computer to this computer. Note that the program in this embodiment includes information for use in processing by an electronic computer and that which conforms to the program (data having a property of defining the processing of a computer but not being a direct instruction to the computer, etc.).

[0054] In each embodiment, the present device is configured by causing a predetermined program to be executed on a computer, but at least a part of these processing contents may be realized hardware-wise.

[0055] Note that the present invention is not limited to the above-described embodiments. For example, any of the replacements described in Modification Example 1 of the First Embodiment and any of the replacements described in Modification Example 2 of the First Embodiment may be made. Alternatively, any of the replacements described in Modification Example 1 of the Second Embodiment and any of the replacements described in Modification Example 2 of the Second Embodiment may be made. Alternatively, any of the replacements described in Modification Example 1 of the First Embodiment and any of the replacements described in Modification Example 2 of the Second Embodiment may be made. Alternatively, any of the replacements described in Modification Example 1 of the Second Embodiment and any of the replacements described in Modification Example 2 of the First Embodiment may be made.

[0056] Further, instead of the DNN, other models such as a Hidden Markov Model (HMM) may be used.

[0057] In addition, the above-described various processes are not only executed in time series according to the description, but may also be executed in parallel or individually according to the processing ability of the device that executes the processes or as necessary. Needless to say, various other changes can be made as appropriate without departing from the spirit of the present invention.

Description of Reference Numerals

[0058] 1,2 Signal processing device

Claims

1. A signal separation unit that applies a separation filter to an observation signal based on source signals emitted from one or more signal sources to obtain a separation signal; A separation filter estimation unit that updates the separation filter based on a frequency covariance matrix that approximates the frequency correlation of the source signals by a linear combination of a first matrix that is a diagonal matrix having, as diagonal elements, the squares of the separation signals in each frequency band representing the estimated power of the source signals in each frequency band based on the observation signal or the separation signal, and a second matrix that represents a rank-1 frequency covariance corresponding to the estimated waveform of the source signals based on the observation signal or the separation signal, and the observation signal; comprising; The second matrix is a matrix having, as elements, the separation signals in each frequency band. A signal processing device.

2. The signal processing device according to Claim 1, wherein the frequency covariance matrix is a linear combination of a first weighted matrix obtained by weighting the first matrix with a non-zero first weight and a second weighted matrix obtained by weighting the second matrix with a non-zero second weight. A signal processing device.

3. The signal processing device according to Claim 2, wherein The frequency covariance matrix is R jn = α 1 D jn + α 2 (v jn → )(v jn → ) H is represented by D jn is the first matrix, and d 1jn 2 , …, d Ijn 2 is an I×I diagonal matrix diag{d 1jn 2 , …, d Ijn 2} with diagonal elements, and (v jn → )(v jn → ) H is the second matrix, and v jn → is v 1jn ,…,v Ijn is an I-dimensional vector with elements i = 1, …, I are indices representing predetermined frequency intervals, j = 1, …, J are indices representing predetermined time intervals, n = 1, …, N are indices representing each of the signal sources, I, J, N are positive integers, α 1 is the first weight, α 2 is the second weight, β H represents the Hermitian transpose of β, the observed signal is a time-frequency domain signal, and the separated signal is a time-frequency domain signal. β ijn A signal processing apparatus representing elements corresponding to the frequency interval, the time interval, and the signal source, which are represented by i, j, and n.

4. A frequency covariance estimation unit that obtains and outputs a frequency covariance matrix that approximates the frequency correlation of the source signals by a linear combination of a first matrix that is a diagonal matrix having, as diagonal elements, the squares of the separation signals in each frequency band representing the estimated power of the source signals in each frequency band based on an observation signal based on source signals emitted from one or more signal sources or a separation signal obtained by applying a separation filter to the observation signal, and a second matrix that represents a rank-1 frequency covariance corresponding to the estimated waveform of the source signals based on the observation signal or the separation signal; The second matrix is a matrix having, as elements, the separation signals in each frequency band. A signal processing device.

5. A signal separation step of applying a separation filter to an observation signal based on source signals emitted from one or more signal sources to obtain a separation signal; A separation filter estimation step of updating the separation filter based on a frequency covariance matrix that approximates the frequency correlation of the source signals by a linear combination of a first matrix that is a diagonal matrix having, as diagonal elements, the squares of the separation signals in each frequency band representing the estimated power of the source signals in each frequency band based on the observation signal or the separation signal, and a second matrix that represents a rank-1 frequency covariance corresponding to the estimated waveform of the source signals based on the observation signal or the separation signal, and the observation signal; comprising; A signal processing method, wherein the second matrix is a matrix having, as elements, the separated signals in respective frequency intervals. **Claim 6** A frequency covariance estimation for obtaining and outputting a frequency covariance matrix that approximates the frequency correlation of a source signal by a linear combination of a first matrix, which is a diagonal matrix having, as diagonal elements, the powers of the separated signals in respective frequency intervals representing the estimated power of the source signal in the respective frequency intervals based on an observation signal based on source signals emitted from one or more signal sources or the separated signals obtained by applying a separation filter to the observation signal, and a second matrix representing a rank-1 frequency covariance corresponding to the estimated waveform of the source signal based on the observation signal or the separated signals, A signal processing method, wherein the second matrix is a matrix having, as elements, the separated signals in respective frequency intervals. **Claim 7** A program for causing a computer to function as the signal processing apparatus according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Signal analysis device, signal analysis method, and signal analysis program

    WO2019163487A1