Source signal separation device, source signal separation method, program, and, information storage medium
By estimating spatial correlation matrices that depend on time and frequency, the source signal separation device improves the precision and versatility of blind source signal separation techniques.
Patent Information
- Application Number
- JP2025086005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-02-22
AI Technical Summary
Existing blind source signal separation techniques, such as FastFCA and FastMNMF, do not fully utilize the dependencies of the spatial correlation matrix on time and frequency, leading to suboptimal accuracy and limited applicability in various situations.
A source signal separation device that estimates spatial correlation matrices dependent on time and frequency, using complex spectrograms and covariance matrices to perform joint diagonalization, allowing for high-precision separation of source signals.
This approach enhances the accuracy and adaptability of blind source signal separation, enabling more precise recovery of source signals from mixed observations.
Smart Images

Figure 2025114867000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a source signal separation device, a source signal separation method, a program, and an information recording medium that are suitable for high-precision blind source separation (BSS). [Background technology]
[0002] As a front-end for acoustic event and speech recognition in real environments, research is being conducted on source signal separation technology, which separates multiple observed signals recorded by a microphone array into multiple source signals. More generally, this technology can be called a signal source separation technology that not only separates acoustic signals but also recovers source signals from observed signals obtained by observing a mixture of various signals at multiple locations in parallel.
[0003] (FastFCA, FastMNMF) Non-Patent Documents 1 and 2 propose a technique for converting observed signals of multiple channels into a complex spectrogram consisting of T time frames, F frequency bins, and M channels, and then performing source signal separation using a spatial correlation matrix of size M that represents the correlation between the M channels. In this technique, a spatial correlation matrix G for each source signal n emitted from each source n among N sources (hereinafter, the "nth source" will be referred to as "source n" where appropriate, and the "nth source signal" will be referred to as "source signal n" where appropriate) and each frequency f of the F frequency bins (hereinafter, the "fth frequency bin" will be referred to as "frequency f" where appropriate) is calculated. nf may differ for each frequency f, but the matrix Q is common to all source signals. f In order to speed up the processing, the constraint that simultaneous diagonalization is possible by f depends on the frequency f but not on the source n. Also, the spatial correlation matrix G nf and matrix Q fis assumed to be time-invariant and does not depend on each time t of the T time frames (hereinafter, the "t-th time frame" will be referred to as "time t" where appropriate).
[0004]
number
[0005]
number
[0006]
number
[0007] In general, for a matrix Q, Q -1 is the inverse matrix of Q, Q H is the Hermitian transpose of Q (a matrix obtained by transposing a matrix and taking the complex conjugate of each element; also known as the "adjoint matrix" or "conjugate transpose"), Q -H is the Hermitian transpose of the inverse of Q.
[0008] Also, the diagonal matrix in which each element of an arbitrary vector a is arranged on the diagonal component is denoted as Diag(a). Also, if we have N matrices A1, ..., A N The block diagonal matrix obtained by arranging the above in order on the diagonal is denoted as Diag(A·). Depending on the context, these may also be denoted as [a] or [A·].
[0009] On the other hand, for any N vectors v1, …, v N The matrix in which the elements of the vector are arranged perpendicular to the direction of the vector elements is [v1, …, v N ] is written as follows.
[0010] Also, for any N scalar values s1, …, s NThe row vectors [s1, …, s N ], the column vector in which these are arranged vertically is [s1, …, s N ] T This becomes:
[0011]
number
[0012] In addition, "diagonalizing matrix Y by matrix U" means that matrix UYU H This means that the diagonal matrix Diag(y) for the vector y is UYU H = Diag(y)
[0013]
number
[0014] Now, g nf is an M-dimensional non-negative vector that can be different for each source n and frequency f.
[0015] Here, in the complex spectrogram of the observed signal, we have a complex vector x of size M at time t and frequency f. tf The spatial correlation matrix of G nf is.
[0016] On the other hand, the complex vector x tf niQ f The new complex vector Q obtained by applying f x tf The covariance matrix of Q f G nf Q f H becomes.
[0017] Therefore, the constraint on the simultaneous diagonalization disclosed in Non-Patent Document 2 is "x tf Q f This means that in the transformed space, each channel is independent. [Prior art documents] [Non-patent literature]
[0018] [Non-Patent Document 1] Nobutaka Ito, Shoko Araki, and Tomohiro Nakatani, "FastFCA: A Joint Diagonalization Based Fast Algorithm for Audio Source Separation Using A Full-Rank Spatial Covariance Model", arXiv:1805.06572v1, https: / / arxiv.org / abs / 1805.06572v1, May 17, 2018 [Non-patent document 2] Kouhei Sekiguchi, Aditya Arie Nugraha, Yoshiaki Bando, and Kazuyoshi Yoshii, "Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices", arXiv:1903.03237v1, https: / / arxiv.org / abs / 1903.03237v1, March 8, 2019 Summary of the Invention [Problem to be solved by the invention]
[0019] In the techniques disclosed in Non-Patent Document 1 and Non-Patent Document 2, the spatial correlation matrix G that depends on the source signal n and the frequency f is nf is a vector g that does not depend on time t (time-invariant) and can take different values for each source signal n and frequency f. nf are diagonal components.
[0020] However, there is a demand for further improvement in accuracy by selecting a more appropriate representation for the spatial correlation matrix.
[0021] There is also a demand for the simultaneous diagonalization technique to be applicable to a variety of situations.
[0022] The present invention has been made to solve the above problems, and has as its object to provide a source signal separation device, a source signal separation method, a program, and an information recording medium that are suitable for high-precision blind source signal separation. [Means for solving the problem]
[0023] The source signal separation device according to the present invention comprises: an acquisition unit that acquires observed signals observed at a plurality of positions; a transformer for transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y1, ..., Y N an estimation unit for estimating The complex spectrogram and the N covariance matrices Y1, ..., Y N and a separation unit that separates N source signals by Equipped with The estimation unit a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimate is made by restricting the number of possible
[0024]
number
[0025]
number
[0026]
number
[0027] According to the present invention, it is possible to provide a source signal separation device, a source signal separation method, a program, and an information recording medium that are suitable for high-precision blind source signal separation. [Brief explanation of the drawings]
[0028] [Figure 1] 1 is an explanatory diagram showing a schematic configuration of a source signal separation device according to an embodiment of the present invention;
[0029] [Figure 2] 4 is a flowchart showing a control flow of a separation process executed by a source signal separation device according to an embodiment of the present invention.
[0030] [Figure 3] 10 is a graph showing changes in cost function values when repeated calculations are performed in a source signal separation device.
[0031] [Figure 4] 10 is a flowchart showing the flow of NF-IVA control executed by the source signal separation device according to the embodiment of the present invention.
[0032] [Figure 5] FIG. 10 is an explanatory diagram showing a neural network configuration of a flow block. DETAILED DESCRIPTION OF THE INVENTION
[0033] The following describes embodiments of the present invention. Note that these embodiments are for illustrative purposes only and do not limit the scope of the present invention. Therefore, those skilled in the art can adopt embodiments in which each or all of the elements of the present embodiments are replaced with equivalents. Furthermore, elements described in each example can be omitted as appropriate depending on the application. In this way, all embodiments constructed in accordance with the principles of the present invention are included in the scope of the present invention.
[0034] (composition) The source signal separation device according to this embodiment is typically implemented by a computer running a program. The computer is connected to various output devices and input devices, and transmits and receives information to and from these devices.
[0035] A program executed by a computer can be distributed or sold by a server to which the computer is connected for communication, or it can be recorded on a non-transitory information recording medium such as a CD-ROM (Compact Disk Read Only Memory), flash memory, or EEPROM (Electrically Erasable Programmable ROM), and then the information recording medium can be distributed, sold, etc.
[0036] The program is installed on a non-transitory information recording medium such as a hard disk, solid-state drive, flash memory, EEPROM, etc., possessed by the computer. The information processing device of this embodiment is then realized by the computer. Generally, the computer's central processing unit (CPU) reads the program from the information recording medium into random access memory (RAM) under the control of the computer's operating system (OS), and then interprets and executes the code contained in the program. However, in an architecture in which the information recording medium can be mapped within a memory space accessible by the CPU, explicit loading of the program into RAM may not be necessary. Various pieces of information required during program execution can be temporarily stored in RAM.
[0037] Furthermore, as mentioned above, it is desirable for the computer to be equipped with a GPU (Graphics Processing Unit) for performing various image processing calculations at high speed. By using a GPU and a library such as PyTorch, it becomes possible to utilize the learning functions and source signal separation functions in various artificial intelligence processes under the control of the CPU.
[0038] It should be noted that the information processing device of this embodiment may be configured using a dedicated electronic circuit rather than a general-purpose computer. In this embodiment, the program may be used as a resource for generating wiring diagrams, timing charts, and the like for the electronic circuit. In this embodiment, an electronic circuit that meets the specifications defined in the program is configured using an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the electronic circuit functions as a dedicated device that performs the functions defined in the program, thereby realizing the information processing device of this embodiment.
[0039] In the following, for ease of understanding, the source signal separation device will be described assuming an embodiment in which the source signal separation device is realized by a computer executing a program. Fig. 1 is an explanatory diagram showing a schematic configuration of a source signal separation device according to an embodiment of the present invention.
[0040] As shown in the figure, a source signal separating device 101 according to this embodiment includes an acquiring unit 102 , a converting unit 103 , an estimating unit 104 , and a separating unit 105 .
[0041] In this embodiment, each unit is realized by a computer executing a program.
[0042] 2 is a flowchart showing the flow of control of the separation process executed by the source signal separation device according to the embodiment of the present invention. The following description will be made with reference to this figure.
[0043] First, the acquisition unit 102 of the source signal separating device 101 acquires observed signals observed at a plurality of positions (step S201).
[0044] The source signal separating device 101 according to this embodiment performs blind source signal separation of M observed signals into N source signals. When applied to sound source separation, the original sound is obtained from mixed sounds (observed signals) obtained by observing a situation in which sounds (source signals) emitted from N sound sources are mixed using M microphones. In this case, the obtaining unit 102 obtains the obtained M mixed sounds as observed signals.
[0045] The observation signals obtained here may be obtained in real time from a plurality of microphones connected to a computer via an A / D converter, or may be pre-recorded data.
[0046] Next, the conversion unit 103 of the source signal separation device 101 converts the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels (step S202).T×F×M By serializing in order of channel frequency, the vector s ∈ C TFM can be obtained. Considering the sampling rate and sampling time when obtaining the observation signal, it is assumed that T < FM, F < TM, M < T without loss of generality. The short-time Fourier transform (STFT) can be used for this conversion.
[0047]
Number
[0048] Here, the image of the source signal n (the multi-channel complex spectrogram of the source signal n) S n ∈ C T×F×M is serialized into a vector s n ∈ C TFM which is assumed to follow a multivariate complex Gaussian distribution based on the covariance matrix Y n ∈ S TFM of all elements.
[0049] Then, the estimator 104 of the source signal separation device 101 estimates N covariance matrices Y1,..., Y for each of the N sources from the complex spectrogram S (steps S203 - S206). N (steps S203 - S206).
[0050] The estimation performed here is, after initializing various parameters (step S203), until the various parameters converge (step S204), calculating a cost function based on the log-likelihood, etc. (step S205), and updating the various parameters by a predetermined algorithm such as the steepest descent method so as to minimize the cost (step S206), and repeating the process (returning to step S204).
[0051]
Number
[0052]
number
[0053] The algorithm, which calculates the log-likelihood and cost function based on the current parameters and updates the various parameters to optimize them, repeating this process until the various parameters converge, can be considered a type of learning. Therefore, applying libraries such as PyTorch, which are prepared for neural network technology, to these calculations enables high-speed processing.
[0054] If the algorithm has a convergence guarantee, the convergence of the parameters can be determined by whether or not a preset number of iterations have been performed. Alternatively, similar to known convergence determination techniques, the amount or ratio of the parameter that has changed due to the update can be calculated, and convergence can be determined when this amount or ratio falls below a threshold.
[0055] Now, additivity of complex spectrograms s = Σ n=1 N s n ; Y = Σ n=1 N Y n Assuming that, the reproducibility of the complex Gaussian distribution gives us the following equation:
[0056]
number
[0057]
number
[0058]
number
[0059] where observation X = ssH Then the log-likelihood and cost functions are:
[0060]
number
[0061]
number
[0062]
number
[0063] Covariance matrix Y1, …, Y N Once the covariance parameter Y is estimated, the covariance parameter Y is determined. Therefore, by using a multivariate Wiener filter, the observed image s can be obtained from the vector s relating to the multi-channel complex spectrogram S based on the observed signal. n can be inferred a posteriori.
[0064] That is, the separation unit 106 of the source signal separation device 101 receives a complex spectrogram S and N covariance matrices Y1, ..., Y n From this, N source signals are separated (step S207), and this process ends.
[0065] Here, using the vector s that serializes the complex spectrogram S, the distribution of the image of the source signal n can be expressed as follows:
[0066]
number
[0067] The expected value E[s n ] is as follows:
[0068]
number
[0069] Therefore, E[s n ], the source signal n is separated.
[0070] Below, we will explain three technologies for signal separation (FastCTF, FastMNMF, and FastMCTF).
[0071] (Decomposition by FastCTF) The above formulation corresponds to approximating the observation X with the covariance parameter Y.
[0072] Here, in CTF (Correlation Tensor Decomposition), we have a positive semidefinite matrix X∈S TF is a set of positive semidefinite matrices H1, …, H according to the basis k=1, …, K. K ∈S T and the group of positive semidefinite matrices W1, …, W K ∈S F In the following, in order to extend the CTF, which is originally a single-channel method, to all channels, we use an M-row M-column (size M) identity matrix I M We take the Kronecker product with
[0073]
number
[0074]
number
[0075] To speed up parameter estimation by this approximation, FastCTF uses H k and W k Assume that is jointly diagonalizable.
[0076]
number
[0077]
number
[0078]
number
[0079]
number
[0080] where h k ∈R T and w k ∈R F are non-negative vectors, which are the time basis vector and the frequency basis vector for basis k, respectively.
[0081] In this case, the covariance parameter Y can be written as follows:
[0082]
number
[0083]
number
[0084] (Decomposition by FastMNMF) Images of sound source n n The spectrum s of the channel axis at time t and frequency f ntf ∈C M and then assume that these follow independent complex Gaussian distributions.
[0085] Then, the covariance matrix Y n ∈S TFM is a block diagonal matrix. The covariance matrix Y n The (t,f)th block diagonal element of Y ntf ∈SM Let's say.
[0086]
number
[0087] The spatial correlation matrix of source signal n at time t and frequency f is G ntf and assume that the power spectral density over time and frequency has a low-rank structure.
[0088]
number
[0089] where h nk = [h nk1 , …, h nkT ] T ∈R T is the time basis vector for source signal n, and w nk = [w nk1 , …, w nkF ] T ∈R F is the frequency basis vector for source signal n.
[0090]
number
[0091]
number
[0092] In MNMF (non-negative matrix factorization), the spatial correlation matrix is assumed to be time-invariant. For any time t, G ntf =G nf Therefore, the following holds:
[0093]
number
[0094] Then, X = ss H The covariance parameter Y that approximates can be written as:
[0095]
number
[0096] Here, Diag(G n· ) is G n1 , …, G nF The operator with a dot inside a circle means the Hadamard product (element product), which is the product of elements at the same position in the matrix. Note that the Hadamard product will be represented by "◎" as appropriate due to restrictions on the types of characters that can be used in the text of this specification. M is an M-by-M matrix with all elements equal to 1.
[0097] To perform maximum likelihood estimation of parameters under this assumption, the MM (Minorization Maximization) algorithm can be adopted, and the computational complexity in this case is O(TFM 3 )
[0098] In FastMNMF, the time-invariant spatial correlation matrix G at each frequency f is n1 , …, G nF the time-invariant matrix Q f In order to speed up the maximum likelihood estimation of parameters, we limit the diagonalization to those that can be performed simultaneously.
[0099]
number
[0100]
number
[0101]
number
[0102] g n is a non-negative vector, and Q f Under the assumption that is a regular matrix, the covariance parameter Y can be rewritten as follows:
[0103]
number
[0104] Here, the matrix U is given as follows:
[0105]
number
[0106] Then, from the reproducibility of the complex Gaussian distribution, the following holds for the matrix U and the vector s based on the complex spectrogram S of the observed signal:
[0107]
number
[0108] Here, Us is the slice s of vector s along the channel axis at time t and frequency f. tf ∈C M A, Q f ∈C M×M Q converted by f s tf ∈C M Q f s tf Us has TFM elements, all of which are independent, and Q f will have a function similar to that of a separation filter.
[0109] For the FastMNMF formulated as above, the maximum likelihood estimation algorithm of parameters disclosed in Non-Patent Document 2 can be applied, and the order of computational complexity is O(TFM 2 )
[0110] (Decomposition by MCTF) CTF performs source signal separation based on a time-frequency covariance model. On the other hand, MNMF performs source signal separation based on a spatial covariance model. Therefore, below we will explain a method that aims to further improve the accuracy of signal separation by combining these two methods and using MCTF decomposition.
[0111] To take into account the time-frequency variance in the decomposition in MNMF, we first remove the time-invariant restriction and perform time-varying joint diagonalization.
[0112]
number
[0113]
number
[0114]
number
[0115] In FastMCTF, which is based on FastCTF and FastMNMF, the matrix U can be formulated as follows:
[0116]
number
[0117] The matrix U formulated here reduces to a CTF if M = 1 (the source signal n = 1, ..., N can be replaced with a basis k = 1, ..., K), and reduces to a FastMNMF if R and P are identity matrices.
[0118] In this case, the covariance matrix Y for the source signal n is n is diagonalized as follows:
[0119]
number
[0120]
number
[0121] λ n is the power spectral density vector for source signal n, and g n is the weight vector for the source signal n. nk is the time basis vector for source signal n and basis k, and w nk are the frequency basis vectors for source signal n and basis k.
[0122] In FastMCTF, the complex spectrogram (third-order complex tensor) S∈C of the observed signal is T×F×M A new tensor V∈C is created by applying matrices P, Q, and R in this order to the frequency, channel, and time axes of T×F×M (Each element is v tfm ) is considered. Regarding the channel axis, a different conversion Q f Then, in the new tensor obtained, TFM elements can be treated independently. Therefore, the power of each element x tfm = |v tfm | 2 A low-rank y tfm All we need to do is consider a problem that approximates
[0123]
number
[0124]
number
[0125] Observation X=ss H The problem of finding the covariance parameter Y that approximates is to estimate the parameters H, W, G, R, P, and Q that minimize the following cost function.
[0126]
number
[0127] This cost function has a similar format to FastCTF, etc., so it is possible to apply an algorithm with a similar convergence guarantee. First, H, W, and G are updated as follows:
[0128]
number
[0129] Next, for R, P, and Q, the following updates are performed in the same way as in the independent vector analysis.
[0130]
number
[0131] These updates are performed until convergence of H, W, G, R, P, and Q. Convergence can be determined using various criteria, such as determining convergence after a predetermined number of repeated updates, or determining convergence when the maximum or average value of the rate of change of each matrix element is equal to or less than a predetermined threshold.
[0132] The computational complexity of directly decomposing S while considering the complete covariance in time, frequency, and space is O(T 3 F 3 M 3 ), but by imposing simultaneous diagonalization as a constraint and decomposing each element independently, it is possible to achieve high-speed execution in O(TFM(T+F+M)).
[0133] As mentioned above, FastMCTF combines FastCTF and FastMNMF to impose the time-varying simultaneous diagonalization constraint, but if R and P are fixed to identity matrices in FastMCTF and excluded from the update targets, it is possible to realize FastMNMF with the time-invariant simultaneous diagonalization constraint. This is equivalent to imposing the following constraint:
[0134]
number
[0135]
number
[0136]
number
[0137] (FastMCTF evaluation experiment) Below, we explain the results of evaluation experiments conducted to confirm the behavior of FastMCTF.
[0138] In this evaluation experiment, we used a four-channel sound mixture containing male and female voices from the Spatialize WSJ0-mix dataset (N=2, M=4). The sound sources and microphones were randomly positioned, and no reverberation was included. A complex spectrogram of the sound mixture was obtained using STFT with a window width of 512 and a shift length of 256 (T=295, F=257).
[0139] FastMCTF, which uses all covariance matrices, was compared with FastMNMF, which uses spatial covariance matrices, and FastMPSDTF-T and FastMPSDTF-F, which use time-frequency covariance matrices in addition to spatial covariance matrices.
[0140] Specifically, P = I F and R = I T After updating H, W, G, and Q 100 times while keeping P fixed (corresponding to FastMNMF), H, W, G, and Q were updated 50 times every time P or R was updated 10 times.
[0141] Figure 3 is a graph showing the change in the cost function value when repeated calculations are performed in a source signal separation device. For all methods, there was no significant difference in the SDR (Signal to Distortion Ratio), and the cost function value decreased monotonically. Because the decrease was large when updating P and R, the change appears to be stepped. Ultimately, the cost function value for FastMCTF was the smallest. Therefore, FastMCTF and FastMNMF provided a foothold for improving the accuracy of source signal separation.
[0142] In the above experiment, simple repeated calculations were performed, but in FastMCTF and other methods, the model has a high degree of freedom, so it is possible to avoid falling into poor local solutions by adopting methods such as simulated annealing.
[0143] (NF-IVA) The matrix Q in the above method f , Q tf can be considered as a matrix that transforms the observed signal into the source signal. The constraint of simultaneous diagonalization imposed by the above method is f , Q tf This is equivalent to considering the observation signals transformed in the above manner to be independent.
[0144] Therefore, the matrix Q of the above method can be calculated by the normalizing flow (NF) that can be realized by neural networks. f , Q tf It is possible to calculate the equivalent of and perform Independent Vector Analysis (IVA).
[0145] The NF-IVA according to this embodiment calculates the STFT coefficient vector x of the observed signal at frequency f and time t under the condition M=N. ft = [x 1ft , …, x Mft ] T ∈C M the STFT coefficient vector y ft = [y 1ft , …, y Nft ] T ∈C N The decomposition matrix W to be converted to ft is calculated using a neural network. In the following explanation, the order of the frequency and time subscripts has been changed from "tf" to "ft." Also, although both m and n are used as indexes below, the constraint M=N is imposed, so m and n can be interchanged.
[0146]
number
[0147]
number
[0148] where x ft is a rearrangement of the complex spectrogram S of the observed signal in the above embodiment, and is substantially equivalent to this.
[0149] In the NF-IVA according to this embodiment, L flow blocks are used to ft From y ftCalculate.
[0150]
number
[0151] where K = 2L+1 is the total number of flows.
[0152] W k',f is the time-invariant projection matrix.
[0153]
number
[0154] W k'',ft =Diag(s k'',ft ) is a time-varying diagonal matrix.
[0155]
number
[0156] s k'',ft is the time-varying diagonal matrix W k'',ft is a vector that arranges the diagonal elements of and combines the coupling function outputs.
[0157]
number
[0158] So overall, W ft functions as a time-varying signal separation filter. ft is the matrix Q of the above method. f , Q tf is equivalent to
[0159]
number
[0160] where the source signal ynt Assume that follows a circularly symmetric Gaussian distribution.
[0161]
number
[0162] Then the log-likelihood function of the source signal and the observation is:
[0163]
number
[0164] where ||·|| F is the Frobenius norm. Also, {v} m is the operation to extract the m-th element of the vector v. Variance σ nt 2 is calculated in maximum likelihood terms as follows:
[0165]
number
[0166] In NF-IVA, the log likelihood ln p(X) is maximized, so L NVP = -ln p(X).
[0167]
number
[0168] 4 is a flowchart showing the flow of NF-IVA control executed by a source signal separating device according to an embodiment of the present invention. The following description will be made with reference to this figure. Note that in the control shown in this figure, STFT and inverse STFT are omitted.
[0169] In the NF-IVA decomposition, first, for any frequency f and any time t, a temporary vector h 0,ftThe observed signal x ft (Step S401).
[0170]
number
[0171] Then, for any odd k' and frequency f, the projection matrix W k',f The initial value of is initialized by, for example, a unit matrix (step S402).
[0172]
number
[0173] Then, for any even k'' and frequency f, the scale estimator Ω a k'',f , Ω b k'',f is initialized by, for example, a uniform distribution (step S403).
[0174]
number
[0175] Then, the repetition is started for the number I (step S404). Note that the number I may be an appropriate constant, or the repetition may be terminated when a predetermined convergence condition is satisfied. Furthermore, instead of repeating for the number I, the repetition may be continued until convergence is achieved.
[0176] That is, repetition for each of the time-frequency bins ft begins (step S405).
[0177] The process starts repeating for each flow block l=1, . . . , L (step S406).
[0178] First, projection is performed (step S407).
[0179]
number
[0180] Next, vector split is performed (step S408).
[0181]
number
[0182] Next, scale estimation is performed (step S409).
[0183]
number
[0184] Then, element-wise scaling is performed (step S410).
[0185]
number
[0186] Furthermore, scale estimation is performed again (step S411).
[0187]
number
[0188] Next, element-wise scaling is performed again (step S412).
[0189]
number
[0190] Then, vector concatination is performed (step S414).
[0191]
number
[0192] After the repetition for each flow block l is completed (step S415), a final projection is performed (step S416). ft is updated (step S417).
[0193]
number
[0194] After the repetition for each time-frequency bin ft is completed (step S418), a negative log-likelihood function is calculated (step S419).
[0195] Then, based on the calculated results, we use the gradient descent method to find all W k',f , Ω a k'',f , Ω b k'',f is updated (step S420).
[0196] After the repetition of I times (step S421), y ft is output as the separation result of the source signal (step S422), and this process ends.
[0197]
number
[0198] 5 is an explanatory diagram showing the configuration of a flow block using a neural network, which shows how the processes in steps S407-S414 are performed by the neural network.
[0199] Each scale estimator Ω a k'',f , Ω b k'',f is a multilayer perceptron (MLP) with one hidden layer.
[0200] Let M' be the result of floor division by 2 of M. a k'',f For , the input dimension is M-M', the hidden dimension is M', and the output dimension is M'. b k'',f For , the input dimension is M', the hidden dimension is M-M', and the output dimension is M-M'.
[0201] For each k', f, Ω a k'',f , Ω b k'',f The total number of weight parameters and bias parameters in 2 +2M pieces.
[0202] Rectified linear units are used as hidden layers, and the hyperbolic tangent function tanh is applied to the output layer as an activation function. The output is scaled so that each element lies between [-1, 1].
[0203] The above NF-IVA can also impose a volume-preserving constraint for normalization. The volume-preserving constraint is achieved by two means.
[0204] First, at the beginning of the forward propagation of the neural network, each matrix W k',fThis allows us to ignore the second term in ln p(X). Orthogonalization is performed as follows:
[0205]
number
[0206] where j is the iteration index and I is the identity matrix.
[0207] To ensure convergence, the following normalization is performed:
[0208]
number
[0209] where ||·||1 is the L1-norm (Manhattan distance).
[0210] If the following condition is met before the maximum number of iterations J = 32 is reached, the iteration is terminated.
[0211]
number
[0212] Second, we add a regularization term based on the L2-norm (Euclidean distance) to make the third term in ln p(X) as close to zero as possible. As a result, the loss function to be minimized during parameter optimization becomes:
[0213]
number
[0214] Below, L NVP The above-mentioned mode that minimizes L is called NF-IVA (NVP). VPThe mode that minimizes is called NF-IVA(VP), where IVA stands for Independent Vector Analysis.
[0215] (NF-IVA evaluation experiment) Below, we will explain the results of experiments to evaluate the performance of NF-IVA.
[0216] In the experiments, we employed the task of separating two source signals from 2, 4, 6, or 8 observed signals (M = 2, 4, 6, or 8), and the task of separating three source signals from 3, 4, 6, or 8 observed signals (M = 3, 4, 6, or 8).
[0217] For the two-source task, 150 observations were randomly selected from the test set wsj0-2mix. For the three-source task, 150 observations were randomly selected from the test set wsj0-3mix. The reverberation time was uniformly distributed from 0.2 to 0.6 seconds. All data were sampled at 16 kHz. STFT coefficients were extracted using a 2048-point Hann window with 75% overlap (F = 1025). The average number of time frames was T ≈ 175.
[0218] For the above parameters, we conducted experiments using NF-IVA(NVP), NF-IVA(VP), AuxIVA, a well-known technique, and IVA-BP as a baseline technique. For all methods, the source signal is assumed to have a circularly symmetric Gaussian distribution with time-varying variance.
[0219] As a post-processing step, we perform a projection-back process, which projects the separated components into the observed mixture space to obtain multi-channel source images and calibrate the energy scale. Note that when M>N, we adopt the N source images with the highest average power.
[0220] For separation performance, we adopted the BSS-Eval metrics consisting of SDR (Signal to Distortion Ratio), ISR (Image to Spatial Distortion Ratio), SIR (Signal to Interference Ratio), and SAR (Signal to Artifacts Ratio).
[0221] Experiments have shown that increasing the number of microphones may not improve performance, likely because more parameters need to be optimized.
[0222] From the viewpoint of SDR, it was found that M=4 is optimal for N=2, and M=6 is optimal for N=3.
[0223] In addition, it was found that AuxIVA converges better when M=2, 4 than when M=6, 8. This is thought to be because the number of parameters to be updated leads to instability in the numerical calculation.
[0224] On the other hand, the backpropagation-based method achieves good convergence regardless of the value of M. However, the computational load is larger than that of AuxIVA.
[0225] NF-IVA(VP) turns out to be significantly slower than NF-IVA(NVP) due to the high cost of matrix orthogonalization.
[0226] The highest SDR values of NF-IVA (NVP) are all close to the highest SDR value of IVA-BP, which means that the flow block is not useful or is not optimized enough.
[0227] On the other hand, the SDR performance of NF-IVA(VP) outperforms other methods in most tasks. NF-IVA(VP) tries to make the second and third terms of ln p(X) zero, but in the real world, these terms never become zero. Nevertheless, the normalization in NF-IVA(VP) seems to contribute to the performance improvement.
[0228] In terms of SDR, NF-IVA(VP) with L=1 performed best for N=2, and NF-IVA(VP) with L=2 or L=4 performed best for N=2.
[0229] For all tasks, L=1 produces good ISR and SAR, suitable for human hearing. The SIR for L=2, 4 is better than the SIR for L=1, which means that interference can be suppressed by flow blocking.
[0230] (generalization) So far, we have explained the FastMNMF, FastMCTF, and NF-IVA technologies for blind source signal separation.
[0231] These techniques involve time-varying transformations (R; W k'',ft ) and the time-invariant transformation (P, Q; W k',f ) can be thought of as separating the source signal from the observed signal.
[0232] Therefore, the source signal separation device according to this embodiment: Obtaining observed signals observed at multiple locations; Maximum likelihood estimation of time-varying and time-invariant transformations for separating the observed signals into source signals; Separating the source signal from the observed signal using the estimated time-varying and time-invariant transformations It can be generalized as such.
[0233] (summary) As described above, the source signal separation device according to this embodiment has the following features: an acquisition unit that acquires observed signals observed at a plurality of positions; a transformer for transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y1, ..., Y N an estimation unit for estimating The complex spectrogram and the N covariance matrices Y1, ..., Y N and a separation unit that separates N source signals by A source signal separation device comprising: The estimation unit a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimate is made by restricting the number of possible
[0234] In addition, in the source signal separation device according to this embodiment, The estimation unit estimates the N covariance matrices Y1, ..., Y N The respective covariance matrices Y n of, A complex matrix R of size T, A complex matrix P of size F, identity matrix I of size M M and, time-invariant complex matrix Q f are arranged for each frequency f = 1, …, F in F frequency bins to produce F time-invariant complex matrices Q1, …, Q F The block diagonal matrix Diag(Q · )and, The Kronecker product of matrices (×) and The matrix defined by U = R(×)(Diag(Q · )(P(×)I M )) Therefore, A power spectral density vector λ for each source n of the N sources n Diagonal matrix Diag(λ n )and, time-invariant weight vector g n Diagonal matrix Diag(g n )and, Using the diagonalization UY n U H = Diag(λ n ) (×) Diag(g n ) The above estimation is made by restricting the It can be configured as follows.
[0235] In addition, in the source signal separation device according to this embodiment, The estimation unit calculates a power spectral density vector λ for each source n of the N sources. n is a non-negative matrix factorization in K bases. Diag(λ n ) = Σ k=1 K Diag(h nk ) (×) Diag(w nk ) Therefore, a respective source n of the N sources; For each base k of the K bases, Time basis vector hnk and, Frequency basis vector w nk and, The above estimation is made by limiting the decomposition to It can be configured as follows.
[0236] In addition, in the source signal separation device according to this embodiment, The estimation unit calculates the N covariance matrices Y1, ..., Y N Estimate the It can be configured as follows.
[0237] In addition, in the source signal separation device according to this embodiment, The estimation unit calculates the complex matrix Q by the neural network. tf Ask for It can be configured as follows.
[0238] In addition, in the source signal separation device according to this embodiment, The estimation unit estimates the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt At any time t among the T time frames, the time-invariant spatial correlation matrix G nf , a time-invariant complex matrix Q f , and the time-invariant weight vector g n and make the above estimate by restricting It can be configured as follows.
[0239] In addition, in the source signal separation device according to this embodiment, The estimation unit estimates the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt At any time t among the T time frames, the time-invariant spatial correlation matrix G nf , a time-invariant complex matrix Q f , and the time-invariant weight vector g nand the complex matrix R is constrained to be equal to an identity matrix I T and the complex matrix P is converted into a unit matrix I F and the matrix U is U = I T (×)Diag(Q · ) By doing so, the above estimation is made. It can be configured as follows.
[0240] In addition, in the source signal separation device according to this embodiment, The estimation unit calculates the N covariance matrices Y1, ..., Y N Estimate the It can be configured as follows.
[0241] In addition, in the source signal separation device according to this embodiment, The estimation unit uses the neural network to estimate the time-invariant complex matrix Q f Ask for It can be configured as follows.
[0242] In addition, in the source signal separation device according to this embodiment, The neural network is based on normalizing flow It can be configured as follows.
[0243] The source signal separation method according to this embodiment includes the steps of: an acquisition step of acquiring observed signals observed at a plurality of positions; a transformation step of transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y1, ..., Y N an estimation step of estimating The complex spectrogram and the N covariance matrices Y1, ..., Y N and a separation step of separating N source signals by A source signal separation method comprising: In the estimation step, a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimate is made by restricting the number of possible
[0244] The program according to this embodiment executes the following steps: an acquisition unit that acquires observed signals observed at a plurality of positions; a transformer for transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y1, ..., Y N an estimation unit for estimating The complex spectrogram and the N covariance matrices Y1, ..., Y N and a separation unit that separates N source signals by A program that functions as The estimation unit a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G forntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimate is made by restricting the number of possible
[0245] The source signal separation device according to this embodiment includes: an acquisition unit that acquires observed signals observed at a plurality of positions; a transform unit that transforms the observed signal into an STFT coefficient vector; an estimation unit that estimates a projection matrix to be sequentially applied to the STFT coefficient vector using a neural network; a separation unit for separating N source signals using the STFT coefficient vector and the estimated projection matrix; Equipped with.
[0246] In addition, in the source signal separation device according to this embodiment, The neural network is based on normalizing flow It can be configured as follows.
[0247] The source signal separation method according to this embodiment includes the steps of: an acquisition step of acquiring observed signals observed at a plurality of positions; a transformation step of transforming the observed signal into an STFT coefficient vector; an estimation step of estimating projection matrices to be sequentially applied to the STFT coefficient vectors using a neural network; a separation step of separating N source signals using the STFT coefficient vector and the estimated projection matrix; Equipped with.
[0248] The program according to this embodiment executes the following steps: an acquisition unit that acquires observed signals observed at a plurality of positions; a transform unit that transforms the observed signal into an STFT coefficient vector; an estimation unit that estimates a projection matrix to be sequentially applied to the STFT coefficient vector using a neural network; a separation unit for separating N source signals using the STFT coefficient vector and the estimated projection matrix; Function as.
[0249] The source signal separation device according to this embodiment includes: an acquisition unit that acquires observed signals observed at a plurality of positions; an estimation unit that performs maximum likelihood estimation of a time-varying transformation and a time-invariant transformation for separating the observed signals into source signals; a separation unit that separates a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation; Equipped with.
[0250] The source signal separation method according to this embodiment includes the steps of: an acquisition step of acquiring observed signals observed at a plurality of positions; an estimation step of maximum likelihood estimating a time-varying transform and a time-invariant transform for separating the observed signal into source signals; a separation step of separating a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation. Equipped with.
[0251] The program according to this embodiment executes the following steps: an acquisition unit that acquires observed signals observed at a plurality of positions; an estimation unit that performs maximum likelihood estimation of a time-varying transformation and a time-invariant transformation for separating the observed signals into source signals; a separation unit that separates a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation; Function as.
[0252] The program according to this embodiment can be recorded on a non-transitory computer-readable information recording medium and distributed or sold, or can be distributed or sold via a temporary transmission medium such as a computer communication network.
[0253] The present invention allows various embodiments and modifications without departing from the broad spirit and scope of the present invention. Furthermore, the above-described embodiments are intended to explain the present invention and do not limit the scope of the present invention. That is, the scope of the present invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and the meaning of the invention equivalent thereto are considered to be within the scope of the present invention. [Industrial Applicability]
[0254] According to the present invention, it is possible to provide a source signal separation device, a source signal separation method, a program, and an information recording medium that are suitable for high-precision blind source signal separation. [Explanation of symbols]
[0255] 101 Source Signal Separator 102 Acquisition Department 103 Conversion Unit 104 Estimation part 105 Separation section
Claims
1. an acquisition unit that acquires observed signals observed at a plurality of positions; a transformer for transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y for each of the N sources are obtained. 1 , …, Y N an estimation unit for estimating The complex spectrogram and the N covariance matrices Y 1 , …, Y N and a separation unit that separates N source signals by A source signal separation device comprising: The estimation unit a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimation is made by restricting the A source signal separation device comprising:
2. The estimation unit estimates the N covariance matrices Y 1 , …, Y N The covariance matrix Y of each n of, A complex matrix R of size T, A complex matrix P of size F, identity matrix I of size M M and, time-invariant complex matrix Q f are arranged for each frequency f = 1, …, F in F frequency bins, and are then converted into F time-invariant complex matrices Q 1 , …, Q F The block diagonal matrix Diag(Q ・ )and, The Kronecker product of matrices (×) and The matrix defined by U = R(×)(Diag(Q ・ )(P(×)I M )) Therefore, A power spectral density vector λ for each source n of the N sources n Diagonal matrix Diag(λ n )and, time-invariant weight vector g n Diagonal matrix Diag(g n )and, Using the diagonalization UY n U H = Diag(λ n ) (×) Diag(g n ) The above estimation is made by restricting the 2. The source signal separating device according to claim 1.
3. The estimation unit calculates a power spectral density vector λ for each source n of the N sources. n is a non-negative matrix factorization in K bases. Diag(λ n ) = Σ k=1 K God nk ) (×) Diag(w nk ) Therefore, a respective source n of the N sources; For each base k of the K bases, Time basis vector h nk and, Frequency basis vector w nk and, The above estimation is made by limiting it to those that can be decomposed into 3. The source signal separating device according to claim 2.
4. The estimation unit uses a neural network to estimate the N covariance matrices Y 1 , …, Y N Estimate the 2. The source signal separating device according to claim 1.
5. The estimation unit calculates the complex matrix Q by the neural network. tf Ask for 5. The source signal separating device according to claim 4.
6. The estimation unit estimates the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt At any time t among the T time frames, the time-invariant spatial correlation matrix G nf , a time-invariant complex matrix Q f , and the time-invariant weight vector g n and make the above estimate by restricting 2. The source signal separating device according to claim 1.
7. The estimation unit estimates the spatial correlation matrix G ntf , the complex matrix Q tf , and the weight vector g nt At any time t among the T time frames, the time-invariant spatial correlation matrix G nf , a time-invariant complex matrix Q f , and the time-invariant weight vector g n and the complex matrix R is constrained to be equal to an identity matrix I T Then, the complex matrix P is converted into a unit matrix I F and the matrix U is U = I T (×)Diag(Q ・ ) By doing so, the above estimation is made.
3. The source signal separating device according to claim 2.
8. The estimation unit uses a neural network to estimate the N covariance matrices Y 1 , …, Y N Estimate the 7. The source signal separating device according to claim 6.
9. The estimation unit uses the neural network to estimate the time-invariant complex matrix Q f Ask for 9. The source signal separating device according to claim 8.
10. The neural network is based on normalizing flow 10. The source signal separating device according to claim 4, 5, 8 or 9.
11. an acquisition step of acquiring observed signals observed at a plurality of positions; a transformation step of transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y for each of the N sources are obtained. 1 , …, Y N an estimation step of estimating The complex spectrogram and the N covariance matrices Y 1 , …, Y N and a separation step of separating N source signals by A source signal separation method comprising: In the estimation step, a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimation is made by restricting the 1. A source signal separation method comprising:
12. Computer, an acquisition unit that acquires observed signals observed at a plurality of positions; a transformer for transforming the observed signal into a complex spectrogram in T time frames, F frequency bins, and M channels; From the complex spectrogram, N covariance matrices Y for each of the N sources are obtained. 1 , …, Y N an estimation unit for estimating The complex spectrogram and the N covariance matrices Y 1 , …, Y N and a separation unit that separates N source signals by A program that functions as The estimation unit a respective source n of said N sources; a time t for each of the T time frames, and The frequency f of each of the F frequency bins The spatial correlation matrix G for ntf of, A complex matrix Q that may depend on the time t and the frequency f but does not depend on the source n tf Therefore, A weight vector g that may depend on the source n but does not depend on the frequency f nt Diagonal matrix Diag(g nt )fart, Joint Diagonalization Q tf G ntf Q tf H = Diag(g nt ) The above estimation is made by restricting the A program characterized by:
13. an acquisition unit that acquires observed signals observed at a plurality of positions; a transform unit that transforms the observed signal into an STFT coefficient vector; an estimation unit that estimates a projection matrix to be sequentially applied to the STFT coefficient vector using a neural network; a separation unit for separating N source signals using the STFT coefficient vector and the estimated projection matrix; A source signal separation device comprising:
14. The neural network is based on normalizing flow 14. The source signal separating device according to claim 13.
15. an acquisition step of acquiring observed signals observed at a plurality of positions; a transformation step of transforming the observed signal into an STFT coefficient vector; an estimation step of estimating projection matrices to be sequentially applied to the STFT coefficient vectors using a neural network; a separation step of separating N source signals using the STFT coefficient vector and the estimated projection matrix; A source signal separation method comprising:
16. Computer, an acquisition unit that acquires observed signals observed at a plurality of positions; a transform unit that transforms the observed signal into an STFT coefficient vector; an estimation unit that estimates a projection matrix to be sequentially applied to the STFT coefficient vector using a neural network; a separation unit for separating N source signals using the STFT coefficient vector and the estimated projection matrix; A program characterized by functioning as
17. an acquisition unit that acquires observed signals observed at a plurality of positions; an estimation unit that performs maximum likelihood estimation of a time-varying transformation and a time-invariant transformation for separating the observed signals into source signals; a separation unit that separates a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation; A source signal separation device comprising:
18. an acquisition step of acquiring observed signals observed at a plurality of positions; an estimation step of maximum likelihood estimating a time-varying transform and a time-invariant transform for separating the observed signal into source signals; a separation step of separating a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation. A source signal separation method comprising:
19. Computer, an acquisition unit that acquires observed signals observed at a plurality of positions; an estimation unit that performs maximum likelihood estimation of a time-varying transformation and a time-invariant transformation for separating the observed signals into source signals; a separation unit that separates a source signal from the observed signal using the estimated time-varying transformation and the estimated time-invariant transformation; A program characterized by functioning as
20. 20. A computer-readable non-transitory information recording medium on which the program according to any one of claims 12, 16, and 19 is recorded.
Citation Information
Patent Citations
Signal processing system, signal processing method, and signal processing program
JP2018156052A