Source signal separation device, source signal separation method, program, and information recording medium
The source signal separation device employs a complex matrix Q that depends on time and frequency for simultaneous diagonalization, addressing the challenges of high-precision blind source separation by improving matrix selection and diagonalization applicability.
Patent Information
- Application Number
- JP2021025864
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-22
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2041-02-22
AI Technical Summary
Existing source signal separation technologies face challenges in achieving high-precision blind source separation, particularly in selecting appropriate spatial correlation matrices and applying simultaneous diagonalization methods effectively across various situations.
The proposed solution involves a source signal separation device and method that acquires observation signals, converts them into a complex spectrogram, and estimates covariance matrices. It then performs simultaneous diagonalization using a complex matrix Q that depends on time and frequency but not on the source, utilizing a weight vector to form a diagonal matrix for accurate separation.
This approach enables high-precision blind source signal separation by improving the accuracy of spatial correlation matrix selection and making simultaneous diagonalization more applicable, resulting in enhanced signal separation performance.
Smart Images

Figure 0007691092000081 
Figure 0007691092000082 
Figure 0007691092000083
Abstract
Description
Technical Field
[0001] The present invention relates to a source signal separation device, a source signal separation method, a program, and an information recording medium suitable for high-precision blind source separation (BSS).
Background Art
[0002] As a front end for acoustic events and speech recognition in a real environment, research on source signal separation technology for separating a plurality of channels of observed signals recorded by a microphone array into a plurality of source signals has been underway. More generally, this technology can be said to be a signal source separation technology for restoring the source signal from the observed signal obtained by observing in parallel at a plurality of locations a mixture of various signals, not only the separation of acoustic signals.
[0003] (FastFCA, FastMNMF) In Non-Patent Document 1 and Non-Patent Document 2, a plurality of channels of observed signals are converted into a complex spectrogram composed of T time frames, F frequency bins, and M channels, and then source signal separation is performed using a spatial correlation matrix of size M representing the correlation between the M channels. In this technology, for each source signal n emitted from each source n among N sources (hereinafter, the "nth source" will be appropriately described as "source n", and the "nth source signal" will be appropriately described as "source signal n"), and each frequency f of F frequency bins (hereinafter, the "fth frequency bin" will be appropriately described as "frequency f"), the spatial correlation matrix G nf can be different for each frequency f, but is simultaneously diagonalizable by a matrix Q f that is common to all source signals (see Equation (16) in Non-Patent Document 2), thereby aiming to speed up the processing. That is, the matrix Q f depends on the frequency f but does not depend on the source n. Also, the spatial correlation matrix G nf and the matrix Q fis assumed to be time-invariant and does not depend on each time t of T time frames (hereinafter, the "t-th time frame" will be described as "time t" as appropriate).
[0004]
Number
[0005]
Number
[0006]
Number
[0007] In general, for a certain matrix Q, Q -1 is the inverse matrix of Q, Q H is the Hermitian transpose of Q (a matrix obtained by transposing a matrix and taking the complex conjugate of each element, also called the "adjoint matrix" or "conjugate transpose"). Q -H is the Hermitian transpose of the inverse matrix of Q.
[0008] Also, the diagonal matrix with the elements of an arbitrary vector a arranged as diagonal elements is denoted as Diag(a). Also, the block diagonal matrix with arbitrary N matrices A 1 , …, A N arranged in order on the diagonal is denoted as Diag(A·). Depending on the context, these may also be denoted as [a], [A·].
[0009] On the other hand, for arbitrary N vectors v 1 , …, v N arranged in order perpendicular to the direction in which the elements of the vectors are arranged, the matrix is denoted as [v 1 , …, v N .
[0010] Also, for any N scalar values s 1 , …, s N arranged in order horizontally to form a row vector, if it is denoted as [s 1 , …, s N , then the column vector formed by arranging these in order vertically is [s 1 , …, s N T .
[0011]
Mathematics
[0012] In addition, "diagonalizing matrix Y by matrix U" means that the matrix UYU H becomes the diagonal matrix Diag(y) for the vector y. UYU H = Diag(y)
[0013]
Mathematics
[0014] Now, g nf is an M-dimensional non-negative vector, which can be different for each source n and frequency f.
[0015] Here, in the complex spectrogram of the observed signal, the spatial correlation matrix of the complex vector x tf of size M at time t and frequency f is G nf .
[0016] On the other hand, the covariance matrix of the new complex vector Q tf x f obtained by applying Q f to the complex vector x tf is Q f G nf Q f H .
[0017] Therefore, the limitation of simultaneous diagonalization disclosed in Non-Patent Document 2 means that "in the space transformed by tf with Q f , each channel is independent".
Prior Art Documents
Non-Patent Documents
[0018]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0019] Here, in the technologies disclosed in Non-Patent Document 1 and Non-Patent Document 2, the spatial correlation matrix G nf that depends on the source signal n and the frequency f is time-invariant (independent of time t) and can take different values for each source signal n and frequency f, and the vector g nfSimultaneous diagonalization is performed so that they become diagonal components.
[0020] However, there is a desire to further improve the accuracy by selecting a more appropriate expression as the spatial correlation matrix.
[0021] There is also a desire to make the simultaneous diagonalization method applicable to various situations.
[0022] The present invention solves the above problems, and an object thereof is to provide a source signal separation device, a source signal separation method, a program, and an information recording medium suitable for high-precision blind source signal separation.
Means for Solving the Problems
[0023] The source signal separation device according to the present invention An acquisition unit that acquires observation signals observed at a plurality of positions, A conversion unit that converts the observation signal into a complex spectrogram in T time frames, F frequency bins, and M channels, From the complex spectrogram, N covariance matrices Y for each of the N sources 1 , …, Y N An estimation unit that estimates The complex spectrogram and the N covariance matrices Y 1 , …, Y N A separation unit that separates N source signals by Comprising The estimation unit For each source n of the N sources, For each time t of the T time frames, and For each frequency f of the F frequency bins The spatial correlation matrix G ntf For A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf By A weight vector g that may depend on the source n but does not depend on the frequency f nt is used as the diagonal components of a diagonal matrix Diag(g nt ), simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) The estimation is performed while restricting to those for which it is possible.
[0024]
Number
[0025]
Number
[0026]
Number
Advantages of the Invention
[0027] According to the present invention, it is possible to provide a source signal separation device, a source signal separation method, a program, and an information recording medium suitable for high-precision blind source signal separation.
Brief Description of the Drawings
[0028]
Figure 1
[0029]
Figure 2
[0030]
Figure 3
[0031]
Figure 4
[0032]
Figure 5
Embodiments for Carrying Out the Invention
[0033] Embodiments of the present invention will be described below. Note that this embodiment is for explanatory purposes and does not limit the scope of the present invention. Therefore, those skilled in the art can adopt an embodiment in which each element or all elements of this embodiment are replaced with equivalents thereof. Also, the elements described in each example can be appropriately omitted according to the application. Thus, any embodiment configured according to the principle of the present invention is included in the scope of the present invention.
[0034] (Configuration) The source signal separation device according to this embodiment is typically realized by a computer executing a program. The computer is connected to various output devices and input devices and exchanges information with these devices.
[0035] The program executed by the computer can be distributed and sold by a server to which the computer is communicably connected. In addition, after recording on a non-transitory information recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flash memory, or an EEPROM (Electrically Erasable Programmable ROM), it is also possible to distribute and sell the information recording medium.
[0036] The program is installed on a non-transitory information recording medium such as a hard disk, solid state drive, flash memory, EEPROM, etc. that the computer has. Then, the information processing apparatus in the present embodiment is realized by the computer. Generally, the CPU (Central Processing Unit) of the computer reads the program from the information recording medium into the RAM (Random Access Memory) under the management of the computer's OS (Operating System), and then interprets and executes the code included in the program. However, in an architecture where the information recording medium can be mapped within the memory space accessible by the CPU, an explicit loading of the program into the RAM may not be necessary. In addition, various types of information required during the execution of the program can be temporarily recorded in the RAM.
[0037] Furthermore, as described above, it is desirable for the computer to include a GPU (Graphics Processing Unit) to perform various image processing calculations at high speed. By using the GPU and libraries such as PyTorch, it becomes possible to utilize learning functions and source signal separation functions in various artificial intelligence processes under the control of the CPU.
[0038] Note that instead of realizing the information processing apparatus of the present embodiment with a general-purpose computer, it is also possible to configure the information processing apparatus of the present embodiment using a dedicated electronic circuit. In this mode, the program can also be used as a material for generating wiring diagrams, timing charts, etc. of the electronic circuit. In such a mode, an electronic circuit that meets the specifications defined in the program is configured by an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the electronic circuit functions as a dedicated device that performs the functions defined in the program to realize the information processing apparatus of the present embodiment.
[0039] Hereinafter, for ease of understanding, the source signal separation apparatus will be described assuming a mode implemented by a computer executing a program. FIG. 1 is an explanatory diagram showing a schematic configuration of the source signal separation apparatus according to an embodiment of the present invention.
[0040] As shown in this figure, the source signal separation apparatus 101 according to the present embodiment includes an acquisition unit 102, a conversion unit 103, an estimation unit 104, and a separation unit 105.
[0041] In the present embodiment, each unit is implemented by a computer executing a program effectively.
[0042] FIG. 2 is a flowchart showing the control flow of the separation process executed by the source signal separation apparatus according to the embodiment of the present invention. Hereinafter, the description will also refer to this figure.
[0043] First, the acquisition unit 102 of the source signal separation apparatus 101 acquires observation signals observed at a plurality of positions (step S201).
[0044] The source signal separation apparatus 101 according to the present embodiment performs blind source signal separation of M observation signals into N source signals. When applied to sound source separation, the original sound is to be obtained from the mixed sound (observation signal) obtained by observing the situation where sounds (source signals) emitted from N sound sources are mixed with M microphones. In this case, the acquisition unit 102 acquires the obtained M mixed sounds as observation signals.
[0045] The observation signals acquired here may be acquired in real time from a plurality of microphones connected to a computer via an A / D converter, or data recorded in advance may be used.
[0046] Next, the conversion unit 103 of the source signal separation apparatus 101 converts the observation signal into a complex spectrogram in T time frames, F frequency bins, and M channels (step S202). The multi-channel complex spectrogram S of the observation signal ∈ CT×F×M can be obtained by serializing in order of channel frequency, for the vector s ∈ C TFM . Considering the sampling rate and sampling time when obtaining the observation signal, it is assumed that T < FM, F < TM, M < T without loss of generality. The short-time Fourier transform (STFT) can be used for this conversion.
[0047]
Number
[0048] Here, the image of the source signal n (the multi-channel complex spectrogram of the source signal n) S n ∈ C T×F×M serialized into the vector s n ∈ C TFM is assumed to follow a multivariate complex Gaussian distribution based on the covariance matrix Y n ∈ S TFM among all elements.
[0049] Then, the estimation unit 104 of the source signal separation device 101 estimates N covariance matrices Y 1 , …, Y N for each of the N sources from the complex spectrogram S (steps S203 - S206).
[0050] The estimation performed here is, after initializing various parameters (step S203), until the various parameters converge (step S204), calculating a cost function based on the log-likelihood, etc. (step S205), and updating the various parameters by a predetermined algorithm such as the steepest descent method so as to minimize the cost (step S206), and repeating the process (returning to step S204).
[0051]
Number
[0052] [Mathematics]
[0053] Calculating the log-likelihood and cost function based on the current parameters and updating various parameters to optimize them, repeating this until the various parameters converge, can also be considered a type of learning. Therefore, for these calculations, by applying libraries such as PyTorch prepared with neural network technology, high-speed processing becomes possible.
[0054] When there is a convergence guarantee in the algorithm, it is also possible to determine the convergence of the parameters by whether the iteration has been performed a preset number of times. Also, similar to known convergence determination techniques, calculating the amount or ratio of the parameters changed by the update and considering that it has converged when this is less than the threshold value is also acceptable.
[0055] Now, the additivity of the complex spectrogram s = Σ n=1 N s n ; Y = Σ n=1 N Y n Assuming this, the following equation holds due to the reproducibility of the complex Gaussian distribution.
[0056] [Mathematics]
[0057] [Mathematics]
[0058] [Mathematics]
[0059] Here, the observation X = ss H Then, the log-likelihood and the cost function are as follows.
[0060]
Number
[0061]
Number
[0062]
Number
[0063] Covariance matrix Y 1 , …, Y N If is estimated, the covariance parameter Y is determined. Therefore, by using the multivariate Wiener filter, the observation image s can be inferred a posteriori from the vector s related to the multi-channel complex spectrogram S based on the observation signal n .
[0064] That is, the separation unit 106 of the source signal separation device 101 separates N source signals from the complex spectrogram S and N covariance matrices Y 1 , …, Y n and ends this process (step S207).
[0065] Here, when using the vector s obtained by serializing the complex spectrogram S, the distribution of the image of the source signal n is expressed as follows.
[0066]
Number
[0067] The expected value E[s n of the image of the source signal n is as follows.
[0068]
Mathematics
[0069] Therefore, by performing the inverse STFT on E[s n , the source signal n will be separated.
[0070] Next, three techniques (FastCTF, FastMNMF, FastMCTF) for signal separation will be described.
[0071] (Decomposition by FastCTF) The above formulation corresponds to approximating the observation X with the covariance parameter Y.
[0072] Here, in CTF (Correlation Tensor Decomposition), the positive semi-definite matrix X ∈ S TF is decomposed into the sum of Kronecker products of positive semi-definite matrix groups H 1 , …, H K ∈ S T and positive semi-definite matrix groups W 1 , …, W K ∈ S F Below, to extend the originally single-channel method CTF to all channels, the Kronecker product with the identity matrix I M of size M (M rows and M columns) is taken.
[0073]
Mathematics
[0074]
Mathematics
[0075] To accelerate the parameter estimation by this approximation, FastCTF assumes that H k and W k are simultaneously diagonalizable.
[0076] [Number]
[0077] [Number]
[0078] [Number]
[0079] [Number]
[0080] Here, h k ∈ R T and w k ∈ R F are non - negative vectors and are the time - basis vector and frequency - basis vector with respect to the basis k, respectively.
[0081] At this time, the covariance parameter Y can be expressed as follows.
[0082] [Number]
[0083] [Number]
[0084] (Decomposition by FastMNMF) The image s of the sound source n n is divided into the spectrum s of the channel axis at time t and frequency f ntf ∈ C M and it is assumed that these follow an independent complex Gaussian distribution.
[0085] Then, the covariance matrix Y n ∈ S TFMis a block diagonal matrix. The covariance matrix Y n The (t,f)-th block diagonal component of ntf ∈ S M is denoted as
[0086]
Number
[0087] Let G ntf be the spatial correlation matrix of the source signal n at time t and frequency f, and assume that the power spectral density in time and frequency has a low-rank structure.
[0088]
Number
[0089] Here, h nk = [h nk1 , …, h nkT T ∈ R T is the time basis vector of the source signal n, and w nk = [w nk1 , …, w nkF T ∈ R F is the frequency basis vector of the source signal n.
[0090]
Number
[0091]
Number
[0092] In MNMF (Non-Negative Matrix Factorization), the spatial correlation matrix is assumed to be time-invariant. For any time t, since G ntf = G nf the following holds.
[0093] [Mathematics]
[0094] Then, X = ss H The covariance parameter Y approximating [it] can be written as follows.
[0095] [Mathematics]
[0096] Here, Diag(G n· ) is a block diagonal matrix arranging G n1 , …, G nF , and the operator with a dot drawn in the circle means the Hadamard product (element-wise product) taking the product of elements at the same positions of the matrices. Note that the Hadamard product is appropriately marked with "◎" due to the limitation of the character types available in the text of this application. Also, 1 M is an M×M matrix with all elements being 1.
[0097] To perform maximum likelihood estimation of the parameters under this premise, the MM (Minorization Maximization) algorithm can be adopted, and the computational complexity in this case is O(TFM 3 ).
[0098] In FastMNMF, the time-invariant spatial correlation matrices G n1 , …, G nF at each frequency f are restricted to be simultaneously diagonalizable by the time-invariant matrix Q f to speed up the maximum likelihood estimation of the parameters. Simultaneous diagonalization can be expressed as follows.
[0099] [Mathematics]
[0100] [Mathematics]
[0101]
Number
[0102] g n is a non - negative vector, and under the premise that Q f is a regular matrix, the covariance parameter Y can be rewritten as follows.
[0103]
Number
[0104] Here, the matrix U is given as follows.
[0105]
Number
[0106] Then, from the reproductive property of the complex Gaussian distribution, for the matrix U and the vector s based on the complex spectrogram S of the observed signal, the following holds.
[0107]
Number
[0108] Here, Us is "the slice s tf ∈C M in the channel - axis direction at time t and frequency f of the vector s, transformed by Q f ∈C M×M and arranged as Q f s tf ∈C M Q f s tf ". Us has TFM elements, and all of these are independent, and Q f has a function similar to that of a separation filter.
[0109] For the FastMNMF formulated as described above, the estimation algorithm for the maximum likelihood estimation of the parameters disclosed in Non-Patent Document 2 can be applied, and the order of the computational complexity is O(TFM 2 ).
[0110] (Decomposition by MCTF) In CTF, source signal separation is performed based on the covariance model of time and frequency. On the other hand, in MNMF, source signal separation is performed based on the covariance model of space. Therefore, below, a form aiming at further improving the accuracy of signal separation by the decomposition by MCTF combining these will be described.
[0111] In order to be able to consider time-frequency dispersion in the decomposition in MNMF, first, the time-invariant limitation is removed and time-varying simultaneous diagonalization is performed.
[0112]
Number
[0113]
Number
[0114]
Number
[0115] In FastMCTF based on FastCTF and FastMNMF, the matrix U can be formulated as follows.
[0116]
Number
[0117] The matrix U formulated here reduces to CTF if M = 1 (the source signals n = 1, …, N can be rewritten in terms of the basis k = 1, …, K), and to FastMNMF if R and P are identity matrices.
[0118] At this time, the covariance matrix Y for the source signal n n is diagonalized as follows.
[0119]
Number
[0120]
Number
[0121] λ n is the power spectral density vector for the source signal n, and g n is the weight vector for the source signal n. Also, h nk is the time basis vector for the source signal n and the basis k, and w nk is the frequency basis vector for the source signal n and the basis k.
[0122] In FastMCTF, for the frequency axis, channel axis, and time axis of the complex spectrogram (third-order complex tensor) S ∈ C of the observed signal T×F×M a new tensor V ∈ C T×F×M (each element is denoted as v tfm ) is considered by applying the matrices P, Q, and R in this order. For the channel axis, a different transformation Q f is applied for each frequency f. Then, in the resulting new tensor, the TFM elements can be treated independently. Therefore, it suffices to consider the problem of approximating the power x tfm = |v tfm | 2 of each element with a low-rank y tfm .
[0123] [Number]
[0124] [Number]
[0125] Observation X = ss H For the problem of obtaining the covariance parameter Y that approximates [this], it suffices to estimate the parameters H, W, G, R, P, Q that minimize the following cost function.
[0126] [Number]
[0127] Since this cost function has the same form as that of FastCTF, etc., it is possible to apply an algorithm with a similar convergence guarantee. First, for H, W, G, perform the following updates.
[0128] [Number]
[0129] Next, for R, P, Q, perform the following updates in the same way as independent vector analysis.
[0130] [Number]
[0131] Perform these updates until H, W, G, R, P, Q converge. Regarding the convergence determination, various determination criteria can be adopted, such as determining convergence when a predetermined number of iterations of the update have been performed, determining convergence when the maximum value or average value of the change rate of the elements of each matrix is less than or equal to a predetermined threshold value, etc.
[0132] The computational complexity of directly decomposing S while considering the complete covariance with respect to time, frequency, and space is O(T 3 F 3 M 3 ). However, by imposing simultaneous diagonalization as a constraint and performing independent decomposition for each element, fast execution with O(TFM(T + F + M)) becomes possible.
[0133] As described above, FastMCTF combines FastCTF and FastMNMF to impose the constraint of time-varying simultaneous diagonalization. However, if R and P are fixed to the identity matrix and excluded from the update target in FastMCTF, it is possible to realize FastMNMF with the constraint of time-invariant simultaneous diagonalization. This corresponds to imposing the following constraints.
[0134]
Number
[0135]
Number
[0136]
Number
[0137] (Evaluation Experiment of FastMCTF) Below, the results of the evaluation experiment conducted to confirm the behavior of FastMCTF will be described.
[0138] In this evaluation experiment, 4-channel mixed sounds including male and female voices were used in the experiment from the Spatialize WSJ0-mix dataset (N = 2, M = 4). The source and microphone arrangements were random, and reverberation was not included. Using STFT with a window width of 512 and a shift length of 256, the complex spectrogram of the mixed sound was obtained (T = 295, F = 257).
[0139] FastMCTF using all covariance matrices was compared with FastMNMF using spatial covariance matrices, FastMPSDTF-T and FastMPSDTF-F using temporal and frequency covariance matrices in addition to spatial covariance matrices.
[0140] Specifically, with P = I F and R = I T fixed, H, W, G, and Q were updated 100 times (corresponding to FastMNMF), and then H, W, G, and Q were updated 50 times each time P or R was updated 10 times.
[0141] Figure 3 is a graph showing the change in the cost function value during iterative calculations in the source signal separation device. In all methods, there is no significant difference in SDR (Signal to Distortion Ratio), and the cost function value decreases monotonically. Since the decrease is large when updating P or R, the change appears stepwise. Finally, the cost function value of FastMCTF became the minimum. Therefore, clues for improving the accuracy of source signal separation were obtained in FastMCTF and FastMNMF.
[0142] In the above experiment, iterative calculations were simply performed. However, in FastMCTF etc., since the degree of freedom of the model is high, for example, by adopting the simulated annealing method, it is possible to avoid falling into a poor local solution.
[0143] (NF-IVA) Matrix Q in the above method f , Q tf can be considered as a matrix that converts the observed signal into the source signal. And the constraint of simultaneous diagonalization imposed by the above method corresponds to the idea that the observed signals converted by matrix Q f , Q tf are independent.
[0144] Therefore, the matrix Q of the above method can be calculated by normalizing flow (NF) which can be realized by neural networks. f , Q tf It is possible to calculate the equivalent of and perform an Independent Vector Analysis (IVA).
[0145] The NF-IVA according to this embodiment calculates the STFT coefficient vector x of the observed signal at frequency f and time t under the condition M=N. ft = [x 1ft , …, x Mft ] T ∈C M Let us denote the STFT coefficient vector y ft = [y 1ft , …, y Nft ] T ∈C N Decomposition matrix W to be converted to ft is calculated using a neural network. In the following explanation, the order of the frequency and time subscripts is changed from "tf" to "ft." In addition, although both m and n are used as indexes in the following, m and n can be replaced with each other due to the constraint that M=N.
[0146]
number
[0147]
number
[0148] where x ft is a rearrangement of the complex spectrogram S of the observed signal in the above embodiment, and is substantially equivalent to this.
[0149] In the NF-IVA according to this embodiment, L flow blocks are used to ft From ftCalculate it.
[0150]
Number
[0151] Here, K = 2L + 1 is the total number of flows.
[0152] W k',f is a time-invariant projection matrix.
[0153]
Number
[0154] W k'',ft = Diag(s k'',ft ) is a time-varying diagonal matrix.
[0155]
Number
[0156] s k'',ft is a vector formed by arranging the diagonal elements of the time-varying diagonal matrix W k'',ft and is the combined coupling function output.
[0157]
Number
[0158] Therefore, overall, W ft functions as a time-varying signal separation filter. This W ft corresponds to the matrices Q f , Q tf in the above method.
[0159]
Number
[0160] Here, the source signal ynt It is assumed that it follows a circularly symmetric Gaussian distribution.
[0161]
Number
[0162] Then, the logarithmic likelihood functions of the source signal and the observation are as follows.
[0163]
Number
[0164] Here, ||·|| F is the Frobenius norm. Also, {v} m is the operation of extracting the m-th element of the vector v. The variance σ nt 2 is calculated as follows from the maximum likelihood perspective.
[0165]
Number
[0166] In NF-IVA, to maximize the logarithmic likelihood ln p(X), the steepest descent method for minimizing L NVP = -ln p(X) is executed.
[0167]
Number
[0168] Figure 4 is a flowchart showing the control flow of NF-IVA executed by the source signal separation device according to an embodiment of the present invention. Hereinafter, it will be described with reference to this figure. In the control of this figure, the description of STFT and inverse STFT is omitted.
[0169] In the decomposition by NF-IVA, first, for any frequency f and any time t, the temporary vector h 0,ftusing the observation signal x ft for initialization (step S401).
[0170]
Number
[0171] Next, for any odd number k' and frequency f, initialize the initial value of the projection matrix W k',f using, for example, the identity matrix (step S402).
[0172]
Number
[0173] Then, for any even number k'' and frequency f, initialize the scale estimators Ω a k'',f , Ω b k'',f using, for example, a uniform distribution (step S403).
[0174]
Number
[0175] And start the iteration of the number of times I (step S404). Note that the number of times I may adopt an appropriate value as a constant, or the iteration may end when a predetermined convergence condition is satisfied. Also, instead of iterating by the number of times I, it may be continued until convergence.
[0176] That is, start the iteration for each time-frequency bin ft (step S405).
[0177] Start the iteration for each flow block l = 1, …, L (step S406).
[0178] First, perform projection (step S407).
[0179] [Mathematics]
[0180] Next, perform vector split (step S408).
[0181] [Mathematics]
[0182] Then, perform scale estimation (step S409).
[0183] [Mathematics]
[0184] And then, perform element-wise scaling (step S410).
[0185] [Mathematics]
[0186] Furthermore, perform scale estimation again (step S411).
[0187] [Mathematics]
[0188] Then, perform element-wise scaling again (step S412).
[0189] [Mathematics]
[0190] Then, perform vector concatenation (step S414).
[0191]
Number
[0192] When the repetition for each flow block l is finished (step S415), perform the final projection (step S416) and update y ft (step S417).
[0193]
Number
[0194] When the repetition for each time-frequency bin ft is finished (step S418), calculate the negative log-likelihood function (step S419).
[0195] Then, based on the calculated result, update all W k',f , Ω a k'',f , Ω b k'',f (step S420).
[0196] When the repetition for I times is finished (step S421), output y ft as the separation result of the source signal (step S422) and end this process.
[0197]
Number
[0198] FIG. 5 is an explanatory diagram showing a configuration of a neural network of a flow block. This figure shows how the processes according to steps S407 - S414 are performed by the neural network.
[0199] Each scale estimator Ω a k'',f , Ω b k'',f is a multi - layer perceptron (MLP; MultiLayer Perceptron) with one hidden layer.
[0200] Let M' be the result of applying floor division by 2 to M. For Ω a k'',f , the input dimension is M - M', the hidden dimension is M', and the output dimension is M'. For Ω b k'',f , the input dimension is M', the hidden dimension is M - M', and the output dimension is M - M'.
[0201] For each k'', f, for Ω a k'',f , Ω b k'',f the total number of weight parameters and bias parameters is M 2 + 2M.
[0202] The normalized linear unit is used as the hidden layer, and the hyperbolic tangent function tanh is applied as the activation function to the output layer. The output is scaled so that each element is between [-1, 1].
[0203] For the above - mentioned NF - IVA, a volume - preserving constraint for normalization can also be imposed. The volume - preserving constraint is achieved by two means.
[0204] First, at the beginning of starting the forward propagation / forward path of the neural network, for each matrix W k',fOrthogonalize it. This enables the second term of ln p(X) to be ignored. The orthogonalization is performed as follows.
[0205]
Number
[0206] Here, j is the iteration index and I is the identity matrix.
[0207] To ensure convergence, perform the following normalization.
[0208]
Number
[0209] Here, ||·|| 1 is the L1-norm (Manhattan distance).
[0210] Before reaching the maximum number of iterations J = 32, end the iteration if the following condition is satisfied.
[0211]
Number
[0212] Second, add a regularization term based on the L2-norm (Euclidean distance) to make the third term of ln p(X) as close to zero as possible. As a result, the loss function to be minimized during parameter optimization is as follows.
[0213]
Number
[0214] Hereinafter, the above-described manner of minimizing L NVP shall be referred to as NF-IVA (NVP). Also, L VPThe mode that minimizes it will be called NF-IVA(VP). IVA means Independent Vector Analysis.
[0215] (Evaluation Experiment of NF-IVA) Hereinafter, the results of an experiment for evaluating the performance etc. of NF-IVA will be described.
[0216] In the experiment, tasks of separating two source signals from 2, 4, 6, 8 observed signals (M = 2, 4, 6, 8) and separating three source signals from 3, 4, 6, 8 observed signals (M = 3, 4, 6, 8) were adopted.
[0217] In the task related to two source signals, 150 observed signals randomly selected from the test set of wsj0-2mix were adopted. In the task related to three source signals, 150 observed signals randomly selected from the test set of wsj0-3mix were adopted. The reverberation time adopted a uniform distribution from 0.2 seconds to 0.6 seconds. All data was sampled at 16 kHz. The STFT coefficients were extracted using a 2048-point Hann window with 75% overlap (F = 1025). The average number of time frames is T≒175.
[0218] For the above specifications, in addition to NF-IVA(NVP) and NF-IVA(VP), experiments applying the known technology AuxIVA and IVA-BP as a baseline technology were conducted. For all methods, the source signals are assumed to have a circularly symmetric Gaussian distribution with time-varying variance.
[0219] As post-processing, projection-back is executed. Projection-back projects the separated elements into the observed mixed space to obtain a multi-channel source image and calibrates the energy scale. In the case of M>N, the N source images with the highest average power were adopted.
[0220] For the separation performance, the BSS-Eval metrics consisting of SDR (Signal to Distortion Ratio), ISR (Image to Spatial distortion Ratio), SIR (Signal to Interference Ratio), and SAR (Signal to Artifacs Ratio) were adopted.
[0221] According to the experiments, it was found that even if the number of microphones is increased, the performance may not improve. This is presumably because more parameters need to be optimized.
[0222] From the perspective of SDR, the results showed that M = 4 is optimal for N = 2, and M = 6 is optimal for N = 3, respectively.
[0223] In addition, in AuxIVA, it was found that M = 2 and 4 converge better than M = 6 and 8. This is presumably because the number of parameters to be updated leads to numerical instability.
[0224] On the other hand, in the method based on backpropagation, the convergence is good regardless of the value of M. However, the computational load is larger than that of AuxIVA.
[0225] It was found that NF-IVA(VP) is considerably slower than NF-IVA(NVP). This is because the cost of matrix orthogonalization is high.
[0226] The maximum values of SDR of NF-IVA(NVP) are all close to the maximum values of SDR of IVA-BP, which means that the flow block is not useful or the optimization is not sufficient.
[0227] On the one hand, the SDR performance of NF-IVA(VP) outperforms other methods in most tasks. In NF-IVA(VP), an attempt is made to set the second and third terms of ln p(X) to zero, but in the real world, these do not become zero. Nevertheless, the normalization in NF-IVA(VP) seems to contribute to the performance improvement.
[0228] From the perspective of SDR, NF-IVA(VP) with L = 1 for N = 2 achieves the best performance, and for N = 2, NF-IVA(VP) with L = 2 or L = 4 achieves the best performance.
[0229] For all tasks, L = 1 gives good ISR and SAR and is suitable for human listening. The SIR for L = 2, 4 is better than the SIR for L = 1. This means that interference can be suppressed by the flow block.
[0230] (Generalization) So far, the techniques of FastMNMF, FastMCTF, and NF-IVA have been described for blind source signal separation.
[0231] These techniques can be considered to separate the source signals from the observed signals by maximum likelihood estimation of the time-varying transformation (R; W k'',ft ) and the time-invariant transformation (P, Q; W k',f ).
[0232] Therefore, the source signal separation device according to the present embodiment acquires the observed signals observed at a plurality of positions, performs maximum likelihood estimation of the time-varying transformation and the time-invariant transformation for separating the observed signals into source signals, and separates the source signals from the observed signals by the estimated time-varying transformation and the estimated time-invariant transformation and can be generalized as such.
[0233] (Summary) As described above, the source signal separation device according to the present embodiment acquisition unit that acquires observation signals observed at a plurality of positions a conversion unit that converts the observation signal into a complex spectrogram in T time frames, F frequency bins, and M channels from the complex spectrogram, N covariance matrices Y for each of the N sources 1 , …, Y N an estimation unit that estimates a separation unit that separates N source signals using the complex spectrogram and the N covariance matrices Y 1 , …, Y N A source signal separation device comprising: wherein the estimation unit for each source n of the N sources each time t of the T time frames, and each frequency f of the F frequency bins the spatial correlation matrix G for ntf is a complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf by a weight vector g that may depend on the source n but does not depend on the frequency f nt is used as the diagonal component to form a diagonal matrix Diag(g nt ) simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) and the estimation is performed by restricting to those for which it is possible.
[0234] Also, in the source signal separation device according to the present embodiment the estimation unit, for each covariance matrix Y of the N covariance matrices Y 1 , …, Y N each covariance matrix Y n is a complex matrix R of size T a complex matrix P of size F, and an identity matrix I of size M M and, a time-invariant complex matrix Q f are arranged for each of the F frequency bins at frequencies f = 1, …, F, and F time-invariant complex matrices Q 1 , …, Q F to form a block diagonal matrix Diag(Q · ), and the Kronecker product (×) of matrices, and a matrix defined by U = R(×)(Diag(Q · )(P(×)I M )) whereby, the power spectral density vector λ for each source n of the N sources n is used to form a diagonal matrix Diag(λ n ), and a time-invariant weight vector g n is used to form a diagonal matrix Diag(g n ), and to perform diagonalization U Y n U H = Diag(λ n ) (×) Diag(g n ) the estimation is restricted to what is possible, and can be configured as such.
[0235] Also, in the source signal separation device according to the present embodiment, the estimation unit performs non-negative matrix factorization of the power spectral density vector λ for each source n of the N sources n with respect to K bases Diag(λ n ) = Σ k=1 K Diag(h nk ) (×) Diag(w nk ) whereby, each source n of the N sources, and For each of the K bases, basis k, and for the time basis vector h nk and, the frequency basis vector w nk and, limiting the estimation to those that can be decomposed, the estimation can be configured as follows. It can be configured as follows.
[0236] Also, in the source signal separation device according to the present embodiment, the estimation unit estimates the N covariance matrices Y 1 , …, Y N by a neural network. It can be configured as follows.
[0237] Also, in the source signal separation device according to the present embodiment, the estimation unit obtains the complex matrix Q tf by the neural network. It can be configured as follows.
[0238] Also, in the source signal separation device according to the present embodiment, the estimation unit restricts the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt to be equal to the time-invariant spatial correlation matrix G nf , the time-invariant complex matrix Q f , and the time-invariant weight vector g n respectively at any time t of the T time frames, and performs the estimation. It can be configured as follows.
[0239] Also, in the source signal separation device according to the present embodiment, the estimation unit restricts the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt to be equal to the time-invariant spatial correlation matrix G nf, time-invariant complex matrix Q f , and a time-invariant weight vector g n are respectively restricted to be equal, and the complex matrix R is set to the identity matrix I of size T T , and the complex matrix P is set to the identity matrix I of size F F , and the matrix U is U = I T (×)Diag(Q · ) , and thus the estimation can be performed in such a configuration.
[0240] Also, in the source signal separation device according to the present embodiment , the estimation unit performs the estimation of the N covariance matrices Y 1 , …, Y N by means of a neural network in such a configuration.
[0241] Also, in the source signal separation device according to the present embodiment , the estimation unit obtains the time-invariant complex matrix Q f by means of the neural network in such a configuration.
[0242] Also, in the source signal separation device according to the present embodiment , the neural network can be configured based on normalizing flow in such a configuration.
[0243] The source signal separation method according to the present embodiment includes an acquisition step of acquiring observation signals observed at a plurality of positions , a conversion step of converting the observation signals into a complex spectrogram in T time frames, F frequency bins, and M channels , and an estimation step of estimating N covariance matrices Y 1 , …, Y N for each of the N sources from the complex spectrogram The complex spectrogram and the N covariance matrices Y 1 , …, Y N are used to perform a separation process for separating N source signals A source signal separation method comprising: In the estimation process, for each source n of the N sources, for each time t of the T time frames, and for each frequency f of the F frequency bins the spatial correlation matrix G ntf is simultaneously diagonalized tf by a complex matrix Q that depends on the time t and the frequency f but does not depend on the source n nt to a diagonal matrix Diag(g nt ) having the weight vector g Q tf G ntf Q tf H = Diag(g nt ) and the estimation is performed by restricting to those for which this is possible.
[0244] The program according to the present embodiment causes a computer to acquire an observation signal observed at a plurality of positions, convert the observation signal into a complex spectrogram in T time frames, F frequency bins, and M channels, estimate N covariance matrices Y 1 , …, Y N for each of the N sources from the complex spectrogram, and separate N source signals using the complex spectrogram and the N covariance matrices Y 1 , …, Y N and functions as a separation unit A program that causes the computer to function as: The estimation unit each source n of the N sources, each time t of the T time frames, and each frequency f of the F frequency bins for the spatial correlation matrix G ntf is a complex matrix Q that may depend on the time t and the frequency f but not on the source n tf by a weight vector g that may depend on the source n but not on the frequency f nt into a diagonal matrix Diag(g nt ) having the diagonal components simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) the estimation is performed, limited to those for which it is possible.
[0245] The source signal separation device according to the present embodiment includes an acquisition unit that acquires observation signals observed at a plurality of positions, a conversion unit that converts the observation signals into STFT coefficient vectors, an estimation unit that estimates, by a neural network, a projection matrix to be sequentially applied to the STFT coefficient vectors, a separation unit that separates N source signals using the STFT coefficient vectors and the estimated projection matrix and is provided with.
[0246] Also, in the source signal separation device according to the present embodiment, the neural network can be configured to be based on a normalizing flow as described above.
[0247] The source signal separation method according to the present embodiment includes an acquisition step of acquiring observation signals observed at a plurality of positions, a conversion step of converting the observation signals into STFT coefficient vectors, An estimation step of estimating, by a neural network, a projection matrix to be sequentially applied to the STFT coefficient vector A separation step of separating N source signals based on the STFT coefficient vector and the estimated projection matrix comprises
[0248] The program according to this embodiment causes a computer to an acquisition unit that acquires observation signals observed at a plurality of positions a conversion unit that converts the observation signals into STFT coefficient vectors an estimation unit that estimates, by a neural network, a projection matrix to be sequentially applied to the STFT coefficient vector a separation unit that separates N source signals based on the STFT coefficient vector and the estimated projection matrix function as
[0249] The source signal separation device according to this embodiment an acquisition unit that acquires observation signals observed at a plurality of positions an estimation unit that performs maximum likelihood estimation of a time-varying conversion and a time-invariant conversion for separating the observation signals into source signals a separation unit that separates source signals from the observation signals based on the estimated time-varying conversion and the estimated time-invariant conversion comprises
[0250] The source signal separation method according to this embodiment an acquisition step of acquiring observation signals observed at a plurality of positions an estimation step of performing maximum likelihood estimation of a time-varying conversion and a time-invariant conversion for separating the observation signals into source signals a separation step of separating source signals from the observation signals based on the estimated time-varying conversion and the estimated time-invariant conversion comprises
[0251] The program according to this embodiment causes a computer to an acquisition unit that acquires observation signals observed at a plurality of positions An estimation unit that performs maximum likelihood estimation on a time-varying transformation and a time-invariant transformation for separating the observation signal into source signals. A separation unit that separates source signals from the observation signal based on the estimated time-varying transformation and the estimated time-invariant transformation. Function as.
[0252] The program according to this embodiment can be recorded on a non-temporary computer-readable information recording medium and distributed and sold. It can also be distributed and sold via a temporary transmission medium such as a computer communication network.
[0253] The present invention can be implemented in various embodiments and variations without departing from the broad spirit and scope of the present invention. Further, the above-described embodiments are for explaining the present invention and do not limit the scope of the present invention. That is, the scope of the present invention is indicated by the claims rather than the embodiments. And various modifications made within the scope of the claims and within the scope of the meaning of the invention equivalent thereto are considered to be within the scope of the present invention.
Industrial Applicability
[0254] According to the present invention, it is possible to provide a source signal separation device, a source signal separation method, a program, and an information recording medium suitable for high-precision blind source signal separation.
Explanation of Signs
[0255] 101 Source signal separation device 102 Acquisition unit 103 Conversion unit 104 Estimation unit 105 Separation unit
Claims
1. An acquisition unit that acquires observation signals observed at a plurality of positions, A conversion unit that converts the observation signals into a complex spectrogram in T time frames, F frequency bins, and M channels, From the complex spectrogram, an estimator that estimates N covariance matrices Y for each of the N sources 1 , …, Y N ; and an estimator The complex spectrogram and the N covariance matrices Y 1 , …, Y N and a separating unit that separates N source signals based on these A source signal separation device comprising: The estimation unit, For each source n of the N sources, For each time t of the T time frames, and For each frequency f of the F frequency bins Spatial correlation matrix G for ntf is A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf by A weight vector g that may depend on the source n but does not depend on the frequency f nt is used as the diagonal components to form a diagonal matrix Diag(g nt ). Simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) The estimation is performed with the limitation that it is possible, A source signal separation device characterized by this.
2. The estimation unit estimates each of the N covariance matrices Y 1 , …, Y N to the covariance matrix Y n as A complex matrix R of size T, A complex matrix P of size F, Identity matrix I of size M M and Time-invariant complex matrix Q f For each frequency f = 1, …, F of F frequency bins, F time-invariant complex matrices Q 1 , …, Q F consisting of the block diagonal matrix Diag(Q ・ ), and The Kronecker product (×) of matrices, By U = R(×)(Diag(Q ・ )(P(×)I M )) A matrix U defined as By, The power spectral density vector λ for each source n of the N sources n A diagonal matrix Diag(λ n ) having the diagonal components, and Time-invariant weight vector g n A diagonal matrix Diag(g n ) with diagonal elements, and Using, diagonalization U Y n U H = Diag(λ n ) (×) Diag(g n ) The estimation is performed with the limitation that it is possible, The source signal separation device according to claim 1, characterized by this.
3. The estimation unit performs non - negative matrix factorization of the power spectral density vector λ for each source n of the N sources with respect to K bases n Diag(λ n ) = Σ k=1 K Diag(h nk ) (×) Diag(w nk ) By, For each source n of the N sources and For each basis k of the K bases, for Time base vector h nk and Frequency basis vector w nk and The estimation is performed with the limitation that it can be decomposed into, The source signal separation device according to claim 2, characterized by this.
4. The estimation unit estimates the N covariance matrices Y 1 , …, Y N using a neural network The source signal separation device according to claim 1, characterized by this.
5. The estimation unit obtains the complex matrix Q by the neural network tf thereof The source signal separation device according to claim 4, characterized by this.
6. The estimation unit restricts the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt to be equal to the time-invariant spatial correlation matrix G nf , the time-invariant complex matrix Q f , and the time-invariant weight vector g n respectively at any time t of the T time frames, and performs the estimation The source signal separation device according to claim 1, characterized by this.
7. The estimation unit is configured to use the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt such that at any time t of the T time frames, they are respectively equal to the time-invariant spatial correlation matrix G nf , the time-invariant complex matrix Q f , and the time-invariant weight vector g n . The complex matrix R is set to the identity matrix I T of size T, the complex matrix P is set to the identity matrix I F of size F, and the matrix U is U = I T (×)Diag(Q ・ ) By setting, the estimation is performed The source signal separation device according to claim 2, characterized by this.
8. The estimation unit estimates the N covariance matrices Y 1 , …, Y N by using a neural network The source signal separation device according to claim 6, characterized by this.
9. The estimation unit obtains the time-invariant complex matrix Q by the neural network f to obtain The source signal separation device according to claim 8, characterized by this.
10. The neural network is based on normalizing flow The source signal separation device according to any one of claims 4, 5, 8, and 9, characterized by this.
11. An acquisition step of acquiring observation signals observed at a plurality of positions, A conversion step of converting the observation signals into a complex spectrogram in T time frames, F frequency bins, and M channels, An estimation step of estimating N covariance matrices Y for each of the N sources from the complex spectrogram 1 , …, Y N ; The complex spectrogram and the N covariance matrices Y 1 , …, Y N are used to perform a separation process for separating N source signals A source signal separation method comprising: In the estimation step, For each source n of the N sources, For each time t of the T time frames, and For each frequency f of the F frequency bins Spatial correlation matrix G for ntf is A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf is given by A weight vector g that may depend on the source n but does not depend on the frequency f nt is used as the diagonal components to form a diagonal matrix Diag(g nt ), Simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) The estimation is performed with the limitation that it is possible, A source signal separation method characterized by this.
12. A computer, An acquisition unit that acquires observation signals observed at a plurality of positions, A conversion unit that converts the observation signal into a complex spectrogram in T time frames, F frequency bins, and M channels. From the complex spectrogram, an estimator that estimates N covariance matrices Y for each of the N sources 1 , …, Y N ; and an estimation unit The complex spectrogram and the N covariance matrices Y 1 , …, Y N and a separation unit that separates N source signals based on these A program that functions as The estimation unit For each source n of the N sources For each time t of the T time frames, and For each frequency f of the F frequency bins Spatial correlation matrix G for ntf is A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf is given by A weight vector g that may depend on the source n but does not depend on the frequency f nt into a diagonal matrix Diag(g nt ) having the diagonal components Simultaneous diagonalization Q tf G ntf Q tf H = Diag(g nt ) Restrict to those that are possible and perform the estimation A program characterized by this.
13. An acquisition unit that acquires M observation signals observed at a plurality of positions, A conversion unit that converts the observation signal into an STFT coefficient vector, A projection matrix that is sequentially applied to the STFT coefficient vector, The projection matrix applied to the odd-numbered ones is a time-invariant matrix, The projection matrix applied to the even-numbered ones is a time-varying diagonal matrix An estimation unit that estimates the projection matrix by a normalizing flow realized by a neural network, A separation unit that separates N source signals based on the STFT coefficient vector and the estimated projection matrix A source signal separation device comprising
14. M = N The source signal separation device according to claim 13, characterized by this.
15. An acquisition step of acquiring M observation signals observed at a plurality of positions, A conversion step of converting the observation signal into an STFT coefficient vector, A projection matrix that is sequentially applied to the STFT coefficient vector, The projection matrix applied to the odd-numbered ones is a time-invariant matrix, The projection matrix applied to the even-numbered ones is a time-varying diagonal matrix An estimation step of estimating the projection matrix by a normalizing flow realized by a neural network, A separation step of separating N source signals based on the STFT coefficient vector and the estimated projection matrix A source signal separation method characterized by comprising
16. A computer An acquisition unit that acquires M observation signals observed at a plurality of positions, A conversion unit that converts the observation signal into an STFT coefficient vector, A projection matrix that is sequentially applied to the STFT coefficient vector, The projection matrix applied to the odd-numbered ones is a time-invariant matrix, The projection matrix applied to the even-numbered ones is a time-varying diagonal matrix An estimation unit that estimates the projection matrix by a normalizing flow realized by a neural network, A separation unit that separates N source signals based on the STFT coefficient vector and the estimated projection matrix A program characterized by causing it to function as
17. A computer-readable non-transitory information recording medium on which the program according to any one of Claims 12 to 16 is recorded.