Signal separation device, signal separation method, and program

mfFCA enhances signal separation by accounting for time delays through a spatial correlation matrix, addressing the performance limitations of FCA and FCAd in reverberant environments, achieving improved separation accuracy and stability.

JP7740369B2Active Publication Date: 2025-09-17NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023565699
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-09-17
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

Conventional blind signal separation methods, such as FCA and FCAd, face challenges in achieving sufficient separation performance, particularly when the reverberation time is long, leading to a decrease in performance with increasing iterations of the parameter update algorithm.

Method used

The proposed method, multi-frame FCA (mfFCA), extends the conventional FCA model to account for time delays by incorporating a spatial correlation matrix that represents the correlation between sound source components across multiple time frames, using an EM algorithm to optimize parameters and separate target signals.

Benefits of technology

mfFCA effectively separates source signals with high accuracy even when there are time delays, improving separation performance compared to conventional methods by approximately 2 dB on average, especially in environments with long reverberation times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740369000039
    Figure 0007740369000039
  • Figure 0007740369000040
    Figure 0007740369000040
  • Figure 0007740369000041
    Figure 0007740369000041
Patent Text Reader

Abstract

A signal separating device according to one embodiment includes: a creation unit that uses a first observation signal vector representing an observation signal in which target signals from a plurality of signal sources are mixed together, and a time delay set with elements consisting of the time delay until some of the target signals are observed, to create a second observation signal vector which takes the time delay into consideration to expand the first observation signal vector; an optimization unit that is configured to optimize, through a prescribed algorithm, parameters that include a correlation matrix representing transmission characteristics which take into account the time delay of the target signals, and that include the power of the signal source at each time; and a separation unit that is configured to use the optimized parameters and the second observation signal vector to separate the target signals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a signal separation device, a signal separation method, and a program. [Background technology]

[0002] One of the techniques in the field of signal processing is known as blind source separation (BSS), which is a technique for separating a target source signal from a mixture of signals observed by multiple sensors without any information about how the source signals are mixed.

[0003] A method called full-rank spatial covariance analysis (FCA) is known as a blind signal separation method capable of separating signal sources even when the number N is greater than the number M of sensors (Non-Patent Document 1).

[0004] When considering a sound signal as a signal, reverberation generally occurs in spaces such as rooms as the sound reflects off walls. It is known that the separation performance of FCA can deteriorate as the reverberation time increases. This is partly because, when a time waveform sound signal is converted to a frequency waveform using a short-time Fourier transform (STFT), the main part of the reverberation does not fit within the time frame length.

[0005] On the other hand, a method called FCA with delayed source components (hereinafter also referred to as FCAd) is known as a method that also addresses the above-mentioned reverberation by taking into account time-delayed sound source components (Non-Patent Document 2). [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] NQK Duong, E. Vincent, and R. Gribonval, "Underdetermined reverberant audio source separation using a fullrank spatial covariance model," IEEE Trans. Audio, Speech, and Language Processing, vol. 18, no. 7, pp. 1830-1840, Sept. 2010. [Non-patent document 2] M. Togami, "Multi-channel speech source separation and dereverberation with sequential integration of determined and underdetermined models," in Proc. ICASSP, 2020, pp. 231-235. Summary of the Invention [Problem to be solved by the invention]

[0007] However, conventional methods such as FCA and FCAd have difficulty achieving sufficient separation performance. In particular, when the reverberation time is long, conventional methods can experience a decrease in separation performance as the number of iterations of the parameter update algorithm (e.g., the expectation-maximization (EM) algorithm) increases.

[0008] An embodiment of the present invention has been made in view of the above points, and has as its object to accurately separate a source signal from an observed signal even when the source component has a time delay. [Means for solving the problem]

[0009] In order to achieve the above object, a signal separation device according to one embodiment includes: a creation unit configured to create a second observation signal vector by extending a first observation signal vector, the second observation signal vector representing an observation signal in which target signals from multiple signal sources are mixed, and a time delay set whose elements are the time delays until a portion of the target signal is observed, taking the time delays into account; an optimization unit configured to optimize, using a predetermined algorithm, parameters including a correlation matrix representing the transfer characteristics of the target signal taking the time delays into account and the power of the signal sources at each time; and a separation unit configured to separate the target signals using the optimized parameters and the second observation signal vector. [Effects of the Invention]

[0010] Even if the source component has a time delay, the source signal can be separated from the observed signal with high accuracy. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 10 is a diagram illustrating an example of an observed signal vector affected by a sound source component. [Figure 2] FIG. 10 is a diagram illustrating an example of calculation of the Πi operator. [Figure 3] FIG. 10 is a diagram illustrating a calculation example of the BoffDiag operator. [Figure 4] 1 is a diagram illustrating an example of a hardware configuration of a signal separating device according to an embodiment of the present invention. [Figure 5] 1 is a diagram illustrating an example of a functional configuration of a signal separating device according to an embodiment of the present invention. [Figure 6] FIG. 2 is a diagram illustrating an example of a detailed functional configuration of an EM algorithm unit according to the present embodiment. [Figure 7] 10 is a flowchart illustrating a flow of an example of signal separation processing according to the present embodiment. [Figure 8] 1 is a flowchart showing the flow of an example of an EM algorithm according to the present embodiment. [Figure 9] FIG. 1 shows the experimental setup. [Figure 10] FIG. 10 is a diagram showing evaluation results (part 1). [Figure 11] FIG. 10 is a diagram showing evaluation results (part 2). DETAILED DESCRIPTION OF THE INVENTION

[0012] An embodiment of the present invention will be described below. In this embodiment, a signal separating device 10 will be described that targets sound signals and can accurately separate source signals from observed signals even when there is a time delay in the sound source components due to reverberation, etc. However, targeting sound signals is just one example, and any type of signal in which a time delay may occur in the observation of the source components can be targeted.

[0013] <fca> The following describes FCA, which is one of the conventional methods. For details of FCA, please refer to Non-Patent Document 1, etc.

[0014] <Model and Objective Function> Let N>M, the sound source be n=1, ,N, and the sensors be m=1, ,M. In addition, let the time frame when the time waveform sound signal is converted into a frequency waveform by STFT be t∈{1, ,T} and the frequency bin be f∈{1, ,F}. The observed signals of frequency bin f observed by M sensors in time frame t are expressed as an M-dimensional vector x tf =[x 1tf ,···,x Mtf ] Τ ∈C M and is called the observed signal vector.

[0015] In this case, each observed signal vector x tf is the N sound source components c ntf ∈C M Assume that the sum of

[0016]

number

[0017] In addition, each sound source component c ntf is a vector with a mean of 0 and a covariance matrix of C ntf Multivariate Gaussian distribution of

[0018]

number

[0019]

number

[0020] The objective function to be maximized is the sum of the log-likelihoods:

[0021]

number

[0022]

number

[0023] Each sound source component c in the model shown in (1) ntf Since each of the observation signal vectors x tf has a mean vector of 0 and a covariance matrix of X tf Multivariate Gaussian distribution of

[0024]

number

[0025]

number

[0026] <EM algorithm> The objective function (3) can be (locally) maximized using the EM algorithm. For details of the EM algorithm, see, for example, Reference 1.

[0027] In E-Step, the following mean vector μ ntf (c) and the covariance matrix Σ ntf (c) The conditional probability p(c ntf |x tf , θ) is calculated.

[0028]

number

[0029]

number

[0030]

number

[0031] Hereinafter, in the text of the specification, the accent of an accented letter such as the left side of (9) will be written immediately before it. For example, in the text of the specification, the left side of (9) will be written as " ~ C ntf ". Also, in the text of the specification, when calligraphy characters (handwritten-style characters) such as those on the right side of numbers 2 and 6 may be confused with normal characters, they are written with scr immediately before them. For example, in the text of the specification, the calligraphy character N on the right side of number 2 is written as "scrN" because it may be confused with the sound source number N.

[0032] <fcad> One of the conventional methods, FCAd, will be described below. For details of FCAd, please refer to Non-Patent Document 2.

[0033] To take into account reverberation in a space such as a room, the model (FCA model) shown in (1) is extended as follows:

[0034]

number

[0035]

number

[0036] sound source component c ntf (l) represents that the signal output from sound source n in time frame tl is observed in time frame t with a time delay of l. ntf (l) is a vector with a mean of 0 and a covariance matrix of C ntf (l) We assume that the following multivariate Gaussian distribution is

[0037]

number

[0038] Using the model (FCAd model) shown in (10), each observed signal vector x tf has a mean vector of 0 and a covariance matrix of X tf Multivariate Gaussian distribution of

[0039]

number

[0040]

number

[0041] <Proposed method (mfFCA)> The method proposed in this embodiment will be described below. In this embodiment, a method called multi-frame FCA (hereinafter referred to as mfFCA), which is a new extension of FCA, is proposed. mfFCA is a method for multiplying the power s of the same sound source n by ntf The correlation between source components observed in different time frames (e.g., source component c ntf and c n(t+1)f (1) This method takes into account the correlation between the reverberation and the frequency components. This makes it possible to model reverberation across multiple time frames, thereby improving separation performance.

[0042] Since the following description is independent for each frequency bin f, the subscript f representing the frequency bin will be omitted below for simplicity.

[0043] <Sound source components spanning multiple time frames> Power s of source n at time frame t nt Consider the following long sound source component vector, which combines each sound source component (sound source component with no time delay and sound source component with time delay) generated from:

[0044]

number

[0045] And the above sound source component vector - c nt has a mean vector of 0 and a covariance matrix of - C nt Multivariate Gaussian distribution of

[0046]

number

[0047]

number

[0048]

number

[0049] The parameters to be optimized are as follows:

[0050]

number

[0051]

number

[0052]

number

[0053] The above matrix A n (l,l') is the power s of source n at time frame t ntf Two sound source components c originating from and observed in different time frames n(t+l) (l) and c n(t+l') (l') In this way, unlike conventional methods such as FCA and FCAd, mfFCA uses a spatial correlation matrix whose off-diagonal elements represent the correlation between two sound source components that originate from the same sound source in the same time frame and are observed in different time frames. In other words, this spatial correlation matrix can be considered a covariance matrix that represents the transfer characteristics from each sound source to each sensor, including the transfer characteristics between time frames. As will be described later, parameters are optimized using the EM algorithm based on this spatial correlation matrix.

[0054] Probability Model We construct a probabilistic model necessary to consider (13). First, we consider the sound source component vector - c nt Consider which observed signal vectors are affected by scrL={l1, ,l L }, so the sound source component vector - c nt is the time frame t and t+l1,...,t+l L As an example, Fig. 1 shows the case where t=3, scrL={1,2}. In the example shown in Fig. 1, the sound source component vector - c n3 This shows how the observed signal vectors x3, x4, and x5 are affected.

[0055] Therefore, we define the following long observed signal vector:

[0056]

number

[0057]

number

[0058] Next, the sound source components between different sound sources n - c nt are assumed to be independent, and the observed signal vector - x t and{ - c nt Joint probability distribution of |n=1, ,N}

[0059]

number

[0060] The long observed signal vector shown in (17) - x t Assuming that the subvectors that make up { - c nt Given |n=1, ,N}, the conditional probability distribution is

[0061]

number

[0062]

number

[0063] From the above (18) to (23) and the already assumed (14), the observed signal vector - x t and{ - c nt When calculating the joint probability distribution of |n=1, ,N}, the joint probability distribution p( - x t ,{ - c nt |n=1,···,N}|θ) is obtained.

[0064]

number

[0065]

number

[0066] Once the joint probability distribution is obtained, the marginal and conditional probability distributions can be easily obtained, all of which are multivariate Gaussian distributions (see Reference 2).

[0067] The marginal probability distribution can be obtained as follows, where the mean vector is a multivariate Gaussian distribution with covariance matrix (25):

[0068]

number

[0069] In mfFCA, the objective function to be maximized is the sum of the following log-likelihoods:

[0070]

number

[0071] In E-Step, the conditional probability p( - c nt | - x t , θ) is calculated.

[0072]

number

[0073] When calculating the mean vector shown in (29) above, - C nt - X t -1 This part is called a multi-frame multi-channel Wiener filter (Reference 3).

[0074] In the M-Step, the parameter θ is optimized by maximizing a function known as the Q function. If the parameter in the previous iteration of the EM algorithm is θ', the Q function is defined as follows:

[0075]

number

[0076]

number

[0077] Since it is difficult to directly maximize the Q function above with respect to θ, we use S ti We make an approximation by keeping the covariance matrix of (23) fixed at the parameter θ' before updating. This gives us the following update equation:

[0078]

number

[0079]

number

[0080] <Hardware Configuration of Signal Separation Device 10> An example of the hardware configuration of signal separating device 10 according to this embodiment is shown in Fig. 4. As shown in Fig. 4, signal separating device 10 according to this embodiment includes input device 101, display device 102, external I / F 103, communication I / F 104, RAM (Random Access Memory) 105, ROM (Read Only Memory) 106, auxiliary storage device 107, and processor 108. Each of these pieces of hardware is connected to each other via bus 109 so as to be able to communicate with each other.

[0081] The input device 101 is, for example, a keyboard, a mouse, a touch panel, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the signal separating device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.

[0082] The external I / F 103 is an interface with an external device such as a recording medium 103a. The signal separating device 10 can read from and write to the recording medium 103a via the external I / F 103. Examples of the recording medium 103a include a flexible disk, a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0083] The communication I / F 104 is an interface for connecting the signal separating device 10 to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a storage device (storage device) such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The processor 108 is an arithmetic device such as a CPU (Central Processing Unit).

[0084] The signal separating device 10 according to this embodiment has the hardware configuration shown in Fig. 4, and is thereby able to perform the signal separating process described below. Note that the hardware configuration shown in Fig. 4 is merely an example, and the hardware configuration of the signal separating device 10 is not limited to this. For example, the signal separating device 10 may have multiple auxiliary storage devices 107 and multiple processors 108, or may have various other hardware components in addition to the hardware shown in the figure.

[0085] <Functional Configuration of Signal Separation Device 10> An example of the functional configuration of the signal separating device 10 according to this embodiment is shown in Fig. 5. As shown in Fig. 5, the signal separating device 10 according to this embodiment includes an input unit 201, a parameter initialization unit 202, a permutation solving unit 203, an EM algorithm unit 204, a sound source separation unit 205, and an output unit 206. Each of these units is realized by, for example, processing in which one or more programs installed in the signal separating device 10 are executed by the processor 108.

[0086] The input unit 201 receives an observation signal, which is a time waveform, and applies STFT to the signal to generate an M-dimensional (where M is the number of sensors) observation vector x for each time frame t and frequency bin f. tf The input unit 201 also obtains the time lag set scrL={l1, . . . , l L } to obtain the long observed signal vector shown in (17). - x tf Create a.

[0087] The parameter initialization unit 202 initializes the parameter θ shown in (16) (more precisely, the parameter θ expanded to all frequency bins f). Here, the parameter θ expanded to all frequency bins f means that the frequency bin f of the parameter θ shown in (16) is explicitly written and the parameter θ is expanded to all frequency bins f. f When expressed as {θ f |f=1, ,F}. There are various initialization methods, but for example, the method described in Reference 4 can be used. Below, the parameter obtained by expanding (16) to all frequency bins f will be referred to as θ.

[0088] The permutation solving unit 203 solves the sound source component c that is the same in all frequency bins f. ntf The subscript n in the parameter θ is changed so that the subscript n corresponds to the same subscript n. Note that there are various methods for changing the subscript (i.e., permutation solving method), and for example, the method described in Reference 5 may be used.

[0089] The EM algorithm unit 204 optimizes the parameter θ using the EM algorithm. The detailed functional configuration of the EM algorithm unit 204 will be described later. Note that, during the EM algorithm performed by the EM algorithm unit 204, the parameter θ may be passed to the permutation solving unit 203, and the subscript n may be changed (this is indicated by a dashed line in FIG. 5).

[0090] The sound source separation unit 205 separates the observed signal vector - x t Multi-frame multi-channel Wiener filter for - C nt - X t -1 After obtaining the mean vector by applying (35), the separated signal y ntf If the time-delayed sound source component (the second term on the right-hand side of (35)) is not added, a dereverberated separated signal is obtained.

[0091] The output unit 206 outputs the separated signal y ntf Then, the output unit 206 outputs the separated signals to a predetermined output destination.

[0092] Here, a detailed functional configuration of the EM algorithm unit 204 according to this embodiment is shown in Fig. 6. As shown in Fig. 6, the EM algorithm unit 204 according to this embodiment includes a parameter holding unit 211, an observed signal covariance calculation unit 212, an excitation component average covariance calculation unit 213, an excitation component cross product expected value calculation unit 214, a parameter updating unit 215, and a parameter sharing unit 216.

[0093] The parameter storage unit 211 receives the parameter θ and stores it in a memory (for example, the auxiliary storage device 107, etc.).

[0094] The observed signal covariance calculation unit 212 calculates the covariance matrix - X tf Calculate.

[0095] The sound source component average covariance calculation unit 213 calculates the average vector

[0096]

number

[0097]

number

[0098] The sound source component cross product expected value calculation unit 214 calculates the sound source component - c ntf The expected cross product of ~ C ntf Calculate.

[0099] The parameter update unit 215 calculates the parameter θ by (34). ~ A nf and s ntf Update.

[0100] The parameter sharing unit 216 calculates s between a predetermined number (for example, four) of adjacent frequency bins. ntf Specifically, the parameter sharing unit 216 shares s of a predetermined number of adjacent frequency bins. ntf Calculate the average of each s in those frequency bins ntf Then, the parameter sharing unit 216 replaces s included in the parameter θ stored in the memory by the parameter storage unit 211. ntf s after the replacement ntf Rewrite it as:

[0101] <Signal separation processing flow> The flow of signal separation processing according to this embodiment will be described with reference to FIG.

[0102] The input unit 201 receives an observation signal, which is a time waveform (step S101).

[0103] Next, the input unit 201 applies STFT to the observation signal input in step S101 to generate an M-dimensional observation vector x for each time frame t and frequency bin f. tf is obtained (step S102).

[0104] Next, the input unit 201 inputs the time lag set scrL={l1, . . . , l L } to obtain the long observed signal vector shown in (17). - x tf is created (step S103).

[0105] Next, the parameter initialization unit 202 initializes the parameter θ f parameter θ={θ f |f=1, . . . , F} is initialized (step S104).

[0106] Next, the permutation solving unit 203 calculates the same sound source component c ntf are included in the parameter θ so that ~ A nf and s ntf The subscript n is changed (step S105).

[0107] Next, the EM algorithm unit 204 optimizes the parameter θ using the EM algorithm (step S106). Details of this step will be described later.

[0108] Next, the sound source separation unit 205 calculates the observed signal vector - x t Multi-frame multi-channel Wiener filter for - C nt - X t -1 After obtaining the mean vector by applying (35), the separated signal y ntf is obtained (step S107).

[0109] Next, the output unit 206 outputs the separated signal y obtained in step S107. ntf By applying Inverse STFT to the signal, separated signals of time waveforms are obtained (step S108).

[0110] Then, the output unit 206 outputs the separated signals obtained in the above step S108 to any predetermined output destination (step S109), thereby obtaining the desired separated signals.

[0111] <Details of the EM algorithm (step S106)> The flow of the EM algorithm according to this embodiment will be described with reference to FIG.

[0112] The parameter storage unit 211 of the EM algorithm unit 204 receives the parameter θ and stores it in a memory (for example, the auxiliary storage device 107) (step S201).

[0113] The observed signal covariance calculation unit 212 of the EM algorithm unit 204 calculates the covariance matrix - X tf is calculated (step S202).

[0114] Next, the sound source component average covariance calculation unit 213 of the EM algorithm unit 204 calculates the average vector and the covariance matrix according to (29) (step S203).

[0115] Next, the sound source component cross product expectation value calculation unit 214 of the EM algorithm unit 204 calculates the sound source component - c ntf The expected cross product of ~ C ntf is calculated (step S204).

[0116] Next, the parameter update unit 215 of the EM algorithm unit 204 updates the parameter θ by (34). ~ A nf and s ntf is updated (step S205).

[0117] Next, the parameter sharing unit 216 of the EM algorithm unit 204 calculates the s of a predetermined number (for example, four) of adjacent frequency bins. ntf Calculate the average of each s in those frequency bins ntf is replaced with the mean, and the s stored in memory is ntf is rewritten (step S206).

[0118] Then, the EM algorithm unit 204 determines whether a predetermined termination condition is satisfied (step S207). If the EM algorithm unit 204 determines that the termination condition is satisfied, it terminates the EM algorithm, and if not, it returns to step S202. Here, examples of the termination condition include when steps S202 to S206 are repeated a predetermined number of times, or when the improvement in the sum of log-likelihoods (27), which is the objective function, is equal to or less than a predetermined amount.

[0119] <Experiment and its evaluation> Experiments were conducted to evaluate mfFCA.

[0120] The experimental setup was as shown in Fig. 9, with N = 4 and M = 3. That is, microphones 301 to 303 were installed near the center of the room as sensors, and loudspeakers 401 to 404 were installed as sound sources at positions of 70°, 150°, 245°, and 315° around them on a 120 cm circumference. The room size was 4.45 x 3.55 x 2.5 m, and the microphones 301 to 303 and loudspeakers 401 to 404 were installed at a height of 120 cm.

[0121] The reverberation time in the room was varied from 130 ms to 450 ms, and the observed signal at each reverberation time was a mixture of an impulse response and a 6-second speech (English). The sampling frequency was 8 kHz, the STFT analysis window length was 128 ms, and the shift length was 32 ms. Therefore, T = 201 and F = 513.

[0122] The indicator used to evaluate separation performance was signal-to-distortion ratios (SDRs) (Reference 6).

[0123] The evaluation results are shown in Figure 10. In Figure 10, "FCAd{2}" represents the FCAd with scrL={2}, and "mfFCA{2}" represents the mfFCA with scrL={2}. Similarly, "FCAd{2,4}" represents the FCAd with scrL={2,4}, and "mfFCA{2,4}" represents the mfFCA with scrL={2,4}. As Figure 10 shows, except for the shortest reverberation time (130 ms), mfFCA achieves an average performance improvement of approximately 2 dB compared to FCAd.

[0124] The convergence process is shown in Figure 11. With the conventional methods (FCA and FCAd), performance deteriorates as the number of iterations of the parameter update algorithm increases, especially when the reverberation time is long. On the other hand, with mfFCA, performance improves as the number of iterations increases.

[0125] As described above, the proposed method (mfFCA) achieves higher separation performance than the conventional methods (FCA and FCAd), and the separation performance can be improved as the number of iterations of the algorithm for parameter update increases.

[0126] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0127] [References] Reference 1: AP Dempster, NM Laird, and DB Rubin, "Maximum likelihood from incomplete data via the EM algorithm," Journal of the Royal Statistical Society. Series B (Methodological), vol. 39, no. 1, pp. 1-22, 1977. Reference 2: CM Bishop, Pattern Recognition and Machine Learning, Springer, 2006. Reference 3: Z.-Q. Wang, H. Erdogan, S. Wisdom, K. Wilson, D. Raj, S. Watanabe, Z. Chen, and JR Hershey, "Sequential multiframe neural beamforming for speech separation and enhancement," in 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021, pp. 905-911. Reference 4: H. Sawada, R. Ikeshita, N. Ito, and T. Nakatani, "Computational acceleration and smart initialization of full-rank spatial covariance analysis," in Proc. EUSIPCO, 2019, pp. 1-5. Reference 5: H. Sawada, S. Araki, and S. Makino, "Underdetermined convolutive blind source separation via frequency bin-wise clustering and permutation alignment," IEEE Trans. Audio, Speech, and Language Processing, vol. 19, no. 3, pp. 516-527, Mar. 2011. Reference 6: E. Vincent, S. Araki, F. Theis, G. Nolte, P. Bofill, H. Sawada, A. Ozerov, V. Gowreesunker, D. Lutter, and NQK Duong, "The signal separation evaluation campaign (2007-2010): Achievements and remaining challenges," Signal Processing, vol. 92, no. 8, pp. 1928-1936, Aug. 2012. [Explanation of symbols]

[0128] 10 Signal separation device 101 Force input device 102 Display device 103 External I / F 103a Recording media 104 Communication I / F 105 RAM 106 ROM 107 Memory-Assisted Device 108 processors 109 Bus 201 Input section 202 Parameter initialization section 203 Permutation Solution Unit 204 EM Algorithm Section 205 Sound source separation section 206 Output section 211 Parameter storage unit 212 Observation signal covariance calculation unit 213 Sound source component average covariance calculator 214 Sound source component cross product expectation value calculation unit 215 Parameter Update Unit 216 Parameter Sharing Section< / fcad> < / fca>

Claims

1. a creation unit configured to create a second observation signal vector by extending the first observation signal vector in consideration of the time delays, using a first observation signal vector representing an observation signal in which target signals from a plurality of signal sources are mixed and a time delay set having time delays until a portion of the target signal is observed as elements; an optimization unit configured to optimize parameters including a correlation matrix representing a transfer characteristic of the target signal taking into account the time delay and the power of the signal source at each time, using a predetermined algorithm; a separation unit configured to separate the target signal using the optimized parameters and the second observed signal vector; and where t (1≦t≦T) is time, f (1≦f≦F) is frequency, n (1≦n≦N) is the signal source, and l (1≦l≦L) is an element of the time delay set, and the first observed signal vector is expressed as a sum with respect to n of sound source components cntf(0) representing the target signal from each signal source n, and sound source components cntf(1), ..., cntf(L) representing the target signal from each signal source n taking the time delay into consideration, The correlation matrix is ​​a matrix having M×M (M is the number of sensors observing the observed signals) block matrices as off-diagonal elements that represent the correlation between two sound source components cn(t+l)f(l) and cn(t+l')f(l') (where l≠l', 0≦l, l'≦L) that are generated from the power of signal source n at time t and observed at different times, The creation unit configured to create the second observed signal vector at time t by combining the first observed signal vectors at times t, t+1, ..., t+L; The optimization unit Calculating the mean vector and covariance matrix of a multivariate Gaussian distribution that the conditional probability of obtaining cntf given the second observed signal vector and the parameters follows, where cntf = (cntf(0) , cn(t+1)f(1) , ... , cn(t+L)f(L)); is configured to update the parameters by maximizing a sum of log-likelihoods of conditional probabilities of obtaining the second observed signal vector given the parameters using the mean vector and the covariance matrix; The c ntf follows a multivariate Gaussian distribution with a mean vector of 0 and a covariance matrix that is the product of the power of the signal source at time t and the correlation matrix, The separation unit is The signal separating device is configured to separate the target signal by applying a multi-frame multi-channel Wiener filter to the second observed signal vector to obtain the mean vector.

2. The optimization unit The signal separating device according to claim 1 , configured to maximize the sum of the log-likelihoods by maximizing a Q function represented by an expectation value of the sum of the log-likelihoods.

3. The optimization unit c ntf 3. The signal separating device according to claim 1, wherein the parameter is updated using an expected cross product of:

4. a creation step of creating a second observation signal vector by extending the first observation signal vector in consideration of the time delays, using a first observation signal vector representing an observation signal in which target signals from a plurality of signal sources are mixed and a time delay set having elements representing time delays until a portion of the target signal is observed; an optimization procedure for optimizing parameters including a correlation matrix representing a transfer characteristic of the target signal taking into account the time delay and the power of the signal source at each time, using a predetermined algorithm; a separation step of separating the target signal using the optimized parameters and the second observed signal vector; The computer executes where t (1≦t≦T) is time, f (1≦f≦F) is frequency, n (1≦n≦N) is the signal source, and l (1≦l≦L) is an element of the time delay set, and the first observed signal vector is expressed as a sum with respect to n of sound source components cntf(0) representing the target signal from each signal source n, and sound source components cntf(1), ..., cntf(L) representing the target signal from each signal source n taking the time delay into consideration, The correlation matrix is ​​a matrix having M×M (M is the number of sensors observing the observed signals) block matrices as off-diagonal elements that represent the correlation between two sound source components cn(t+l)f(l) and cn(t+l')f(l') (where l≠l', 0≦l, l'≦L) that are generated from the power of signal source n at time t and observed at different times, The creation procedure is as follows: creating the second observed signal vector at time t by combining the first observed signal vectors at times t, t+1, ..., t+L; The optimization procedure comprises: Calculating the mean vector and covariance matrix of a multivariate Gaussian distribution that the conditional probability of obtaining cntf given the second observed signal vector and the parameters follows, where cntf = (cntf(0) , cn(t+1)f(1) , ... , cn(t+L)f(L)); updating the parameters by maximizing the sum of log-likelihoods of the conditional probabilities of obtaining the second observed signal vector given the parameters using the mean vector and the covariance matrix; The c ntf follows a multivariate Gaussian distribution with a mean vector of 0 and a covariance matrix that is the product of the power of the signal source at time t and the correlation matrix, The separation procedure comprises: The signal separation method further comprises applying a multi-frame multi-channel Wiener filter to the second observed signal vector to obtain the mean vector, thereby separating the target signal.

5. A program that causes a computer to function as the signal separation device according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Signal separation device, signal separation method and program

    JP2019074621A