Sound source separation program, sound source separation method, and sound source separation device

By transforming acoustic signals into the frequency domain and using row elementary transformations to update the separation matrix, the sound source separation program efficiently separates sound sources without inverting matrices, addressing the computational challenges of increasing microphone numbers.

JP7683938B2Active Publication Date: 2025-05-27TOKYO METROPOLITAN PUBLIC UNIVERSITY CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022503752
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2021-02-26
Publication Date
2025-05-27
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Conventional sound source separation methods, such as IP, face increased computational costs as the number of microphones grows, particularly due to the inverse matrix calculation in the separation process.

Method used

A sound source separation program that transforms acoustic signals from the time domain to the frequency domain and performs updates based on row elementary transformations to iteratively minimize an objective function, thereby eliminating the need for inverse matrix calculations.

Benefits of technology

This approach enables fast sound source separation without calculating inverse matrices, significantly reducing computational complexity and improving processing speed as the number of microphones increases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683938000043
    Figure 0007683938000043
  • Figure 0007683938000044
    Figure 0007683938000044
  • Figure 0007683938000045
    Figure 0007683938000045
Patent Text Reader

Abstract

This sound source separation program causes a computer to acquire an acoustic signal, convert the acquired acoustic signal from a time domain to a frequency domain, and perform sound source separation on the acoustic signal converted to the frequency domain by iteratively minimizing an objective function including the quadratic form of a separation vector and the determinant of a separation matrix by performing an update based on an elementary row operation on the separation matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a sound source separation program, a sound source separation method, and a sound source separation device. This application claims priority to U.S. Provisional Application No. 62 / 982,755, filed February 28, 2020, the contents of which are incorporated herein by reference. [Background technology]

[0002] Signals picked up by microphones are often mixed signals that are a mixture of sound source signals and noise signals. For such mixed signals, blind source separation is known as a method for estimating the sound source signal without prior information such as the source profile. In blind source separation, sound sources are separated from the mixed signal using a separation matrix W. Here, when there are N sound sources and M microphones, the separation matrix W is a matrix with N rows and M columns. Here, the observed signal x is expressed as the product of the sound source s before mixing and the mixing matrix A. The separation matrix W is the inverse matrix A of this mixing matrix A. -1 Methods for determining the separation matrix W include, for example, Independent Component Analysis (ICA) and Independent Vector Analysis (IVA).

[0003] Furthermore, in recent years, methods for performing blind source separation have been proposed, such as AuxICA (auxiliary function-based independent component analysis; see, for example, Non-Patent Document 1) and AuxIVA (auxiliary function-based independent vector analysis; see, for example, Non-Patent Document 2), which use auxiliary functions.

[0004] In AuxIVA, the separation matrix is ​​estimated by iteratively minimizing the auxiliary function Q in the following equation (1). Note that in the equations, capital bold letters indicate matrices, lowercase bold variables indicate vectors, and normal lowercase variables indicate scalars.

[0005]

number

[0006] In equation (1), k is the index of the sound source signal, f is the index representing the frequency, and F is the total number of frequencies. f =(w 1f …w Kf ) H is the separation matrix to be estimated, M is the number of sound sources (= the number of microphones), and H is the Hermitian transpose. kf is a semi-positive definite matrix that is calculated in different ways depending on the method, such as ICA or IVA. f Since it is not easy to minimize with respect to, AuxIVA updates the row vectors one by one using the update equations (2) and (3) below.

[0007]

number

[0008]

number

[0009] In addition, in formula (2), V kf is expressed as follows:

[0010]

number

[0011] However, e m is a K-dimensional unit vector whose m-th element is 1 and the other elements are 0. Here, this method is called IP (Iterative Projection). [Prior art documents] [Non-patent literature]

[0012] [Non-Patent Document 1] N. Ono and S. Miyabe, “Auxiliary-function-based independent component analysis for super-Gaussian sources”, Proc. LVA / ICA, vol. 6365, no. 6, pp. 165-172, Sep. 2010. [Non-Patent Document 2] N. Ono, “Stable and fast update rules for independent vector analysis based on auxiliary function technique”, in Proc. IEEE WASPAA, New Paltz, NY, USA, Oct. 2011, pp. 189-192. Summary of the Invention [Problem to be solved by the invention]

[0013] However, in conventional techniques such as IP, there is a problem in that as the number of microphones increases, the computational cost of the inverse matrix calculation in equation (2) increases.

[0014] The present invention has been made in consideration of the above problems, and aims to provide a sound source separation program, a sound source separation method, and a sound source separation device that are capable of separating sound sources quickly without calculating an inverse matrix. [Means for solving the problem]

[0015] In order to achieve the above object, a sound source separation program according to one aspect of the present invention causes a computer to acquire an acoustic signal, transform the acquired acoustic signal from the time domain to the frequency domain, and perform an update based on a row elementary transformation on a separation matrix for the acoustic signal transformed into the frequency domain, thereby iteratively minimizing an objective function including a quadratic form of a separation vector and a determinant of the separation matrix, thereby performing sound source separation.

[0016] In addition, in a sound source separation program according to an aspect of the present invention, the computer is caused to perform updating for each frequency f and between k=1, . . . , M using a transformation formula based on the row elementary transformation of the following formula,

number

[0017] In a sound source separation program according to an aspect of the present invention, the computer is configured to calculate a separation matrix W f The k-th column is determined so as to minimize the function, and the remaining columns are unit matrices. By repeating the above process, the separation matrix W f The user may be made to ask for the following:

[0018] In a sound source separation program according to an aspect of the present invention, the function is expressed as follows:

number

[0019] In order to achieve the above object, a sound source separation method according to one aspect of the present invention includes a sound collection unit having a plurality of microphones that acquires an acoustic signal, a sound source separation unit that transforms the acquired acoustic signal from a time domain to a frequency domain, and the sound source separation unit performs an update based on a row elementary transformation on a separation matrix for the acoustic signal converted to the frequency domain, thereby iteratively minimizing an objective function including a quadratic form of a separation vector and a determinant of the separation matrix, thereby performing sound source separation.

[0020] In order to achieve the above object, a sound source separation device according to one embodiment of the present invention includes a sound collection unit having a plurality of microphones for acquiring sound signals, and a sound source separation unit that transforms the acquired sound signals from the time domain to the frequency domain, and performs an update based on a row elementary transformation on a separation matrix for the sound signals converted to the frequency domain, thereby performing sound source separation by iteratively minimizing an objective function including a quadratic form of a separation vector and a determinant of the separation matrix. Effect of the Invention

[0021] According to the present invention, it is possible to separate sound sources at high speed without calculating an inverse matrix. [Brief description of the drawings]

[0022] [Figure 1] FIG. 1 is a diagram illustrating an overview of blind sound source separation processing. [Diagram 2] FIG. 1 is a diagram illustrating an example of a configuration of a sound source separation device according to an embodiment. [Diagram 3] FIG. 13 is a diagram for explaining an update by row elementary transformation. [Figure 4] FIG. 1 is a diagram for explaining an overview of an auxiliary coefficient method using an auxiliary function. [Diagram 5] FIG. 2 is a diagram illustrating an example of an ISS algorithm for sound source separation according to the embodiment. [Figure 6] FIG. 13 is a diagram illustrating an IP algorithm of a comparative example. [Figure 7] FIG. 11 is a diagram for explaining how to improve update efficiency according to an embodiment. [Figure 8]This is a histogram of the reverberation time of the room used in the simulation. [Figure 9] FIG. 1 shows the SDR after 10M iterations. [Figure 10] FIG. 13 shows the SIR after 10M iterations. [Figure 11] FIG. 13 is a diagram showing calculation times for each repetition. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0024] (overview) First, an overview of the embodiment will be described. Fig. 1 is a diagram showing an overview of blind sound source separation processing. As shown in Fig. 1, in blind sound source separation, a separation filter (separation matrix) W is used to separate a separated sound from a mixed sound. In this embodiment, the calculation of the separation matrix W is performed by updating the rank (number of ranks) of the matrix by 1, instead of updating each row vector. As a result, in this embodiment, it is possible to further increase the speed of blind sound source separation.

[0025] (Example of the configuration of a sound source separation device) Next, a configuration example of a sound source separation device will be described. 2 is a diagram showing an example of the configuration of the sound source separation device 1 according to this embodiment. As shown in FIG. 2, the sound source separation device 1 includes an acquisition unit 11, a sound source separation unit 12, and an output unit 13. The sound source separation unit 12 includes an STFT unit 121 , a separation unit 122 , and an inverse STFT unit 123 .

[0026] (Operation of sound source separation device) Next, the operation of the sound source separation device 1 will be described with reference to FIG. The sound source separation device 1 separates a sound source signal from a mixed signal collected by a microphone 2 (a sound collection unit). The microphone 2 is a microphone array made up of a plurality of microphones.

[0027] The acquiring unit 11 acquires a mixed signal (acoustic signal) output by the microphone 2. The acquiring unit 11 converts the mixed signal from an analog signal to a digital signal, and outputs the converted mixed signal to the sound source separation unit 12.

[0028] The sound source separation unit 12 may be, for example, a personal computer, a CPU (Central Processing Unit), a DSP (Digital Signal Processing Unit), an ASIC (Application Specific Integrated Circuit), or the like.

[0029] The STFT unit 121 transforms the mixed signal output by the acquisition unit 11 from the time domain to the frequency domain by Short-Time Fourier Transform.

[0030] The separation unit 122 performs sound source separation by iteratively minimizing an auxiliary function instead of the separation matrix W for the mixed signals subjected to the short-time Fourier transform. The auxiliary function, processing algorithm, etc. will be described later.

[0031] The inverse STFT unit 123 transforms the frequency domain sound source signal separated by the separation unit 122 from the frequency domain to the time domain by an inverse short-time Fourier transform.

[0032] The output unit 13 outputs the sound source signal separated by the sound source separation unit 12 to an external device (for example, a speaker).

[0033] (Example of signal processing) Next, an example of signal processing in the sound source separation process will be described. In the following example, AuxIVA (Auxiliary Function Type Independent Vector Analysis) will be described as an example, but is not limited thereto. The separation matrix update rule of the embodiment can also be applied to AuxICA (Auxiliary Function Type Independent Component Analysis), ILRMA (Independent Low-Rank MAtrix), etc.

[0034] A mixed sound obtained by mixing K sound sources picked up by M microphones can be expressed as in the following formula (5). Note that in the formulas used in the embodiment, capital bold letters indicate matrices, lowercase bold variables indicate vectors, and normal lowercase variables indicate scalars.

[0035]

number

[0036] In equation (5), x^ m [t] is the signal from the mth microphone, and s^ k [t] is the kth sound source signal, and a^ mk [t] is the impulse response of the microphone signal and the sound source signal. The star represents the convolution operation. In the time-frequency domain, the convolution is multiplication for each frequency, as shown in the following equation (6).

[0037]

number

[0038] In formula (6), x mfn is the short-time Fourier transform of x^m[t], and s kfn s^ k [t] is the short-time Fourier transform of mk [f] is a^ mk is the discrete Fourier transform of [t], where f(=1,…,F) are discrete frequency bins and n(=1,…,N) is the frequency index. Note that equation (6) is a valid approximation when the Fourier transform is sufficiently longer than the impulse response. If the microphone signal and the sound source signal at frequency f are grouped by a vector, the microphone signal can be expressed as a linear mixture of the sound source signal as shown in the following equation (7).

[0039]

number

[0040] In equation (7), A f (A f ) mk =a mkf is the confusion matrix by The purpose of Independent Vector Analysis (IVA) is to calculate the separation matrix W f (=[w 1f ,…,w Mf ] H ) is to be sought.

[0041]

number

[0042] In equation (8), y fn are the separated signals. In IVA, we assume that the information sources are statistically independent and that the distribution of the source signals follows a spherical super-Gaussian distribution (p(s k1n ,…,s kFn )~e -G (√(Σ f s kfn )), where G is, for example, the Laplace function G(r)=r or the Cauchy function G(r)=-log(1+r 2 / v). Under these assumptions, AuxIVA estimates the separation matrix by iteratively minimizing the auxiliary function Q in the following equation (9).

[0043]

number

[0044] In other words, formula (9) is a function consisting of a quadratic form (one term) of the separating vector and a determinant (two terms) of the separating matrix. Note that formula (9) may include other terms. Also, the two terms of formula (9) are not limited to the logarithm of the determinant and may be in other forms. In addition, in formula (9), V kf is expressed as follows:

[0045]

number

[0046] In addition, in equation (10), φ(r) is a nonlinear function that depends on the sound source model, for example, φ(r)=1 / r. kn is expressed as the following equation (11).

[0047]

number

[0048] In conventional AuxIVA and the like, row vectors are updated one by one in sequence using the following equations (12) and (13). In the following description, such a method is called IP (iterative projection).

[0049]

number

[0050]

number

[0051] In such an IP method, as the number of microphones increases, the computational cost of the inverse matrix calculation of equation (12) increases.

[0052] (ISS method of this embodiment) Next, the method of this embodiment will be described. Note that the method of this embodiment is also called ISS (Iterative Source Steering). In this embodiment, instead of updating the separation matrix W for each row vector, the separation matrix W is obtained by performing an update based on row basic transformation as shown in the following equation (14). Note that in the update based on row basic transformation, the process is repeated for each frequency f and between k=1, ..., M.

[0053]

number

[0054] In equation (14), v kf (=(v 1kf ,…,v Mkf ) T (T stands for transpose))) is the unknown vector to be calculated. FIG. 3 is a diagram for explaining the update by row elementary transformation. The region indicated by g101 is a diagram for explaining the update by the ISS method of this embodiment. In this embodiment, the separation matrix W f For (g103), an update is performed using row elementary transformation by multiplying all columns except the k-th column (g103) by the diagonal matrix (g102) from the left. The area indicated by g111 is a diagram for explaining updating by the conventional IP method. In the conventional IP method, the k-th row of the separation matrix (g113) is updated.

[0055] The unknown vector v in equation (14) kf The calculation is done by using the auxiliary function Q(v kf ) to minimize v kf This can be done by finding the

[0056]

number

[0057] If f is omitted in equation (15), the result becomes the following equation (16).

[0058]

number

[0059] In equation (16), V m is expressed as follows:

[0060]

number

[0061] In equations (15) and (16), the asterisk * denotes a complex conjugate. Since the auxiliary function Q can be divided into contributions for each frequency f, the frequency index f will be omitted in the following description. This minimization problem (Equation (18) below) can be solved as shown in Equation (19) below. Note that C in Equation (18) is the set of all complex numbers.

[0062]

number

[0063]

number

[0064] If f is not omitted, the result is the following equation (20).

[0065]

number

[0066] Here, applying the theorem regarding the determinant of a matrix, we obtain the following equation (21).

[0067]

number

[0068] In equation (16), if the constant term is omitted, the auxiliary function Q can be simplified as the following equation (22).

[0069]

number

[0070] v * mk Taking the complex derivative with respect to, we obtain the following equation (23).

[0071]

number

[0072] Setting equation (23) equal to zero gives the desired result. This update equation does not involve a matrix inversion. Also, y kn =w H k x n If we pay attention to the above, the amount of update required is only the following equations (24) and (25). mn ) is a nonlinear function that depends on the sound source model.

[0073]

number

[0074]

number

[0075] If f is not omitted in equations (24) and (25), the following equations (26) and (27) result.

[0076]

number

[0077]

number

[0078] In this embodiment, V m The right-hand side of equations (24) and (25) can be calculated efficiently without calculating all the elements of y n Therefore, in this embodiment, it is sufficient to update the following equation (28).

[0079]

number

[0080] If f is not omitted in equation (28), it becomes the following equation (29).

[0081]

number

[0082] These quantities are needed for m, each of which requires N operations, so the total complexity per update is O(MN). Note that for every k updates, k , and all the demodulation filters need to be changed. kn It is sufficient to update only once per iteration.

[0083] Here, an outline of the auxiliary coefficient method using an auxiliary function will be described. Here, we will explain the minimization problem of function J(θ) (J(θ) → min) as an example. The objective function and auxiliary function are J(θ) = min η Q(θ,η) satisfies the relationship. From this relationship, there exists an auxiliary variable η such that for any auxiliary variable η, the auxiliary function Q(θ,η) ≥ objective function J(θ), and for any parameter θ, the auxiliary variable η satisfies J(θ) = Q(θ,η). In the auxiliary function method, the auxiliary function is alternately minimized for the parameter θ and the auxiliary variable η using the following equations (30) and (31). Note that k is a positive integer representing the iteration order.

[0084]

number

[0085]

number

[0086] Fig. 3 is a diagram for explaining an outline of the auxiliary coefficient method using an auxiliary function, in which the horizontal axis represents the parameter θ. Equation (26) is the current estimate θ = θ (k) The auxiliary function Q(θ,η(k+1) ) is calculated. Equation (27) also calculates the auxiliary function Q(θ,η (k+1) ) is the operation to minimize it. Then, by repeating the iterative process, the parameters are updated and minimized as shown in Figure 3. In this way, the auxiliary function method uses J(θ)=min η This is an algorithm that iteratively minimizes the auxiliary function Q(θ,η) that satisfies the relationship Q(θ,η) (see Reference 1).

[0087] Reference 1: Noritaka Ono, "Optimization Algorithm Using Auxiliary Function Method and Its Application to Acoustic Signal Processing", Acoustical Society of Japan, Journal of the Acoustical Society of Japan, Vol. 68, No. 11, 2012, pp. 566-571

[0088] (Algorithm Description) Next, an example of the ISS algorithm for sound source separation according to this embodiment will be described. FIG. 5 is a diagram illustrating an example of the ISS algorithm for sound source separation according to the present embodiment. fn}, and the separated signals are {y fn}. The following process is repeated from 1 to the maximum value (g201). For all k and n, r kn √(Σ|y kfn |) 2 Substitute. The process is repeated for k from 1 to M (g202). For f, the following process is repeated from 1 to F (g203). v km (except m=k) {(Σ n φ(r mn )y mfn y kfn * ) / (Σ n φ(r mn )|y kfn | 2 )} and v kk to {1-(Σ n φ(r mn )|y kfn | 2 ) (-1 / 2)} and for all n, y fn To (y fn -v k y kfn ) for

[0089] As shown in FIG. 4, in this embodiment, there is no procedure for calculating an inverse matrix and no covariance matrix. The computational complexity is O(FM 2 N) / Repetition.

[0090] (Comparison example: IP algorithm) Here, an example of processing using the above-mentioned IP algorithm will be described. FIG. 6 is a diagram illustrating an IP algorithm of a comparative example. The following process is repeated from 1 to the maximum value (g901). For all k and n, r kn to√(Σ|y kfn |) 2 Substitute. The process is repeated for k from 1 to M (g902). For f, the process is repeated from 1 to F (g903). V km to {1 / N(Σ n φ(r kn )x fn x H fn} and w kf To {(W f V kf ) -1 e k} and w kf to {w kf / √(x H fn V kf w kf )}, for all n, y fn To (x H fn w kf ) for

[0091] (Comparison of computational complexity between IP algorithm and ISS algorithm) Comparing Figures 5 and 6, the IP algorithm uses the separation matrix W fThe cost of calculating the inverse matrix of is O(M 3 ) The cost of computing the covariance matrix is ​​O(M 2 N). The total computational complexity of the IP algorithm is O(FM 3 N) / Repetition.

[0092] FIG. 7 is a diagram for explaining how the update efficiency is improved in this embodiment. AuxIVA-IP updates the rows of the separation matrix W. In contrast, the ISS algorithm of this embodiment updates the columns of the mixing matrix, i.e., A=W -1 The k-th steering vector of is updated. In the update, an approximate inverse matrix is ​​calculated using, for example, the Sherman-Morrison method. -1 The update to is equivalent. For example, the process changes the k-th steering vector by the same amount, as shown in the following equation (32). Note that the mixing matrix A=[a 1 ,…,a M ] follows the steering vector of the sound source.

[0093]

number

[0094] In addition, vector a k +u is the vector {1 / (1-v kk )}a k and the vector {1 / (1-v kk )}a m v m The multiplied vector {v m / (1-v kk )}a m In addition, in the Sherman Morrison formula, W=A -1 Therefore, equation (32) becomes the following equation (33).

[0095]

number

[0096] By identifying it with equation (14), v = Wu(1 + w H k u) -1 It can be seen that: In equation (32), the k-th steering vector is updated by the weighted sum of the steering vectors of the other sources, followed by rescaling. mk is the m-th sound source estimate y m The noise of y k and is expressed as the following equation (34).

[0097]

number

[0098] From the properties of φ(r), φ(r mn ) is small when the mth source is active and large when the mth source is inactive. Therefore, in this embodiment, the kth steering vector is modified by an amount proportional to the mth steering vector. Note that scaling is required in this embodiment to maintain the scale of the signal during the iterative process. This process separates, for example, a first signal g311 from another signal g312.

[0099] Next, an example of a comparison result between the IP algorithm and the ISS algorithm of this embodiment will be described. Separation matrix W in the IP algorithm f The computational complexity of updating the k-th row of the covariance matrix V kf or a linear system. As mentioned above, the computational complexity of the IP algorithm is O(M 3 ), and the computational complexity of the ISS algorithm is O(M 2 N). In the IP algorithm, the Mth row update and the Fth frequency band update are repeated, so the total computational complexity of one iteration is C IP is expressed by the following formula (35), and is at least O(M 4 ).

[0100]

number

[0101] In the ISS algorithm, we calculate equations (19) and (21) for m, k = 1, ..., M at each iteration. kn ,∀ k The computation of n has a complexity of O(FMN) per iteration. Therefore, the overall computational complexity per iteration is C ISS is expressed by the following equation (36).

[0102]

number

[0103] However, the computational complexity of the ISS algorithm is a quadratic function of the number of microphones because it uses a single covariance matrix repeatedly.

[0104] (Verification results) Next, the results of an experiment comparing the IP algorithm of the comparative example with the ISS algorithm of this embodiment will be described.

[0105] First, the experimental environment will be described. The experiment was performed using a Python (registered trademark) package, with the following simulation. We used 100 random rectangular rooms with walls between 6m and 10m wide and ceiling heights ranging from 2.8m to 4.5m. The reverberation time (T 60 ) ranged from 60[ms] to 540[ms]. Figure 8 shows a histogram of the reverberation time of the room used in the simulation. The horizontal axis is the reverberation time RT60 [ms], and the vertical axis is the frequency [kHz].

[0106] The sound sources and microphone array were randomly positioned at least 50 cm apart, away from the walls, and at a height of between 1 and 2 m. The microphone array had 10 microphones, was circular with a radius of 3.2 cm, and had microphone spacing of 2 cm. The distance between the sound source and the center of the microphone array should be at least the critical distance d crit =0.057√(V=T 60 ) [m]. V is the room volume. For the first microphone, we use unit power to normalize the source signal.

[0107] SNR = M / σ 2 n Let us define σ 2 n is the variance of uncorrelated white noise at the microphone. The SNR was fixed at 30 [dB]. Separation was performed for 2, 3, 4, 6, 8, and 10 sound sources.

[0108] The number of sound sources is equal to or less than the number of microphones. The sampling frequency is 16 [kHz], the STFT frame size is 256 [ms], and half overlap. A matching window using a Hamming window is used for analysis and synthesis. In the experiment, the AuxIVA-IP algorithm of the comparative example and the ISS algorithm of this embodiment were each repeated 10M times (M is the number of microphones) for separation. After separation, the scale of the output was projected onto the first microphone and restored.

[0109] Signal-to-distortion ratio (SDR) and signal-to-interference ratio (SIR) were used as evaluation indices. SDR and SIR were measured before and after separation. FIG. 9 is a diagram showing SDR after 10M [times] repetition. FIG. 10 is a diagram showing SIR after 10M [times] repetition. In FIGS. 9 and 10, the horizontal axis is the number of channels, and the vertical axis is the improvement amount [dB]. In FIGS. 9 and 10, the symbol g401 is the result of the AuxIVA-IP algorithm of the comparative example, and the symbol g402 is the result of the ISS algorithm of the present embodiment. As shown in FIGS. 9 and 10, the result of using the ISS algorithm of the present embodiment was equivalent to the result of using the AuxIVA-IP algorithm of the comparative example.

[0110] Next, the results of comparing the time required for the separation calculation will be described. FIG. 11 is a diagram showing the calculation time for each iteration. In FIG. 11, the horizontal axis indicates the channel, and the vertical axis indicates the processing time [ms] for each iteration. In FIG. 11, the reference symbol g451 indicates the result of the AuxIVA-IP algorithm of the comparative example, and the reference symbol g452 indicates the result of the ISS algorithm of this embodiment. In the experiment, 1 to 17 sound sources were confirmed. The simulation was performed on a workstation equipped with a 10-core CPU (Central Processing Unit) with a clock frequency of 3.3 [GHz]. The result in FIG. 11 shows the average execution time for one iteration.

[0111] 11, the ISS algorithm of this embodiment requires less time for calculation as the number of sound sources increases compared to the comparative example. In other words, the ISS algorithm of this embodiment can reduce calculation costs more than the AuxIVA-IP of the comparative example.

[0112] As described above, in this embodiment, iterative source steering for independent vector analysis based on auxiliary function method is introduced to sound source separation. While the comparative example AuxIVA-IP updates the decoded vector alternately, the algorithm in this embodiment performs updates based on row elementary transformations continuously. As a result, in this embodiment, an update rule with low computational complexity without inverse matrix is ​​obtained, which can improve stability and speed, and is an ideal method for important practical implementation. The method of this embodiment updates the steering vector of a certain sound source by an amount proportional to the projection of the residual noise of the other sound source into the sound source subspace. From the simulation results, it was confirmed that the method of this embodiment is efficient for sound source separation, and that the calculation cost can be reduced.

[0113] The above-mentioned voice recognition method, program, and voice recognition device can also be applied to voice recognition systems, remote conferencing systems, web conferencing systems, smart speakers, voice input interfaces for home appliances, hearing aids, robot hearing, and the like.

[0114] In addition, a program for realizing all or part of the functions of the sound source separation unit 12 in the present invention may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into a computer system and executed to perform all or part of the processing performed by the sound source separation unit 12. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage providing environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into a computer system. The term "computer-readable recording medium" also refers to a storage device that holds a program for a certain period of time, such as a volatile memory (RAM) inside a computer system that becomes a server or a client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line.

[0115] The above program may also be transmitted from a computer system in which the program is stored in a storage device or the like to another computer system via a transmission medium, or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The above program may also be for realizing part of the above-mentioned functions. Furthermore, it may be a so-called difference file (difference program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0116] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0117] 1... Sound source separation device, 11... Acquisition unit, 12... Sound source separation unit, 13... Output unit, 121... STFT unit, 122... Separation unit, 123... Inverse STFT unit

Claims

1. Cause a computer to acquire an acoustic signal, convert the acquired acoustic signal from the time domain to the frequency domain, perform an update based on an elementary row operation on the acoustic signal converted to the frequency domain with respect to a separation matrix to iteratively minimize an objective function including a quadratic form of a separation vector and a determinant of the separation matrix, and cause source separation to be performed, A source separation program, wherein the source separation program causes the computer to perform an update for each frequency f and between k = 1,..., M by a conversion formula based on the elementary row operation of the following formula, 【Number 1】 solve an unknown vector vkf = (v1,..., vM)T (T represents transpose, k is the number of the source signal and is an integer from 1 to the number of microphones M, f is an index representing frequency) using the objective function, Wf = (w1f,..., wKf)H is a separation matrix, H is a Hermitian transpose, K is the number of sources, M is the number of microphones that picked up the acoustic signal, K = M, wherein the source separation program causes the computer to for each frequency f, update the separation matrix Wf by multiplying it by a matrix in which the k-th column is determined so as to minimize the objective function and the columns other than the k-th column are identity matrices, and obtain the separation matrix Wf by repeating the update, A source separation program.

2. The objective function is the following formula, 【Number 2】 The separation matrix W f is (w 1f … w Kf ) H , where F is the total number of frequencies, H is the Hermitian transpose, and V kf is the weighted covariance matrix. The source separation program according to Claim 1.

3. A sound collection unit including a plurality of microphones acquires an acoustic signal, a source separation unit converts the acquired acoustic signal from the time domain to the frequency domain, the source separation unit performs an update based on an elementary row operation on the acoustic signal converted to the frequency domain with respect to a separation matrix to iteratively minimize an objective function including a quadratic form of a separation vector and a determinant of the separation matrix, and performs source separation, A source separation method, wherein the source separation unit performs an update for each frequency f and between k = 1,..., M by a conversion formula based on the elementary row operation of the following formula, 【Number 3】 solves an unknown vector vkf = (v1,..., vM)T (T represents transpose, k is the number of the source signal and is an integer from 1 to the number of microphones M, f is an index representing frequency) using the objective function, Wf = (w1f,..., wKf)H is a separation matrix, H is the Hermitian transpose, K is the number of sound sources, M is the number of microphones that picked up the acoustic signal, K = M, wherein the sound source separation unit, for each frequency f, multiplies the separation matrix Wf by a matrix in which the k-th column is determined to minimize the objective function and the columns other than the k-th column are identity matrices, and repeats the update to obtain the separation matrix Wf. A sound source separation method.

4. A sound collection unit including a plurality of microphones for acquiring an acoustic signal, a sound source separation unit that converts the acquired acoustic signal from the time domain to the frequency domain, performs an update based on a row elementary transformation on the acoustic signal converted to the frequency domain, and repeatedly minimizes an objective function including a quadratic form of a separation vector and a determinant of the separation matrix to perform sound source separation, A sound source separation device comprising: wherein the sound source separation unit, for each frequency f and for k = 1,..., M, performs an update according to the following transformation formula based on the row elementary transformation, 【Number 4】 solves an unknown vector vkf = (v1,..., vM)T (T represents transpose, k is the number of the sound source signal and is an integer from 1 to the number of microphones M, f is an index representing frequency) using the objective function, Wf = (w1f,..., wKf)H is a separation matrix, H is the Hermitian transpose, K is the number of sound sources, M is the number of microphones that picked up the acoustic signal, K = M, wherein the sound source separation unit, for each frequency f, multiplies the separation matrix Wf by a matrix in which the k-th column is determined to minimize the objective function and the columns other than the k-th column are identity matrices, and repeats the update to obtain the separation matrix Wf. A sound source separation device.

Citation Information

Patent Citations

  • Signal processing apparatus, method, and program

    JP2014041308A