Source separation method, source separator and speech recognition system
Through frequency transformation and clustering technology, combining spatial cues and acoustic cues, the relative transfer function of the speaker is determined and MIMO beamforming is applied, which solves the problem of separation and enhancement of the voice signal of multiple speakers in the reverberation environment, and achieves efficient voice signal enhancement.
Patent Information
- Application Number
- CN202510264269.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-10
- Publication Date
- 2025-06-03
AI Technical Summary
In reverb environments, it is difficult for the prior art to effectively separate and enhance the voice signals of multiple speakers, especially in the presence of multipath tailing and noise interference.
By receiving or generating sound samples, frequency transformation is performed, and clustering is performed according to the speaker, using spatial cues and acoustic cues to cluster, the relative transfer function of each speaker is determined, and a MIMO beamforming operation is applied to enhance the speech signal.
It realizes the effective separation and enhancement of voice signals of multiple speakers in a reverberation environment, improves the performance of the voice enhancement module, and enhances the voice noise ratio and voice interference ratio.
Smart Images

Figure CN120089153A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of March 10, 2019, the application number of 201980096208.9, and the title of "Speech Enhancement Using Clue-based Clustering". Background Art
[0002] The performance of a speech enhancement module depends on the ability to filter out all interfering signals and leave only the desired speech signal. Interfering signals may be, for example, other speakers, noise from an air conditioner, music, motor noise (e.g., in a car or an airplane), and the noise of a large crowd also known as "cocktail party noise". The performance of a speech enhancement module is usually measured by its ability to improve the speech-to-noise ratio (SNR) or the speech-to-interference ratio (SIR), where the speech-to-noise ratio and the speech-to-interference ratio respectively reflect the ratio of the power of the desired speech signal to the total power of the noise and the ratio of the power of the desired speech signal to the total power of other interfering signals (usually in dB).
[0003] There is an increasing need to perform speech enhancement in a reverberant environment. Summary of the Invention
[0004] A method for speech enhancement can be provided, which may include: receiving or generating a sound sample representing a sound signal received by a microphone array during a given time period; performing a frequency transformation on the sound sample to provide a frequency-transformed sample; clustering the frequency-transformed sample according to speakers to provide speaker-related clusters, where the clustering may be based on (i) spatial clues related to the received sound signal and (ii) acoustic clues related to the speakers; determining a relative transfer function for each of the speakers to provide speaker-related relative transfer functions; applying a multiple-input multiple-output (MIMO) beamforming operation on the speaker-related relative transfer functions to provide a beamformed signal; and performing an inverse frequency transformation on the beamformed signal to provide a speech signal.
[0005] The method may include generating acoustic clues related to the speakers.
[0006] The generation of the acoustic clues may include: searching for keywords in the sound sample; and
[0007] extracting the acoustic clues from the keywords.
[0008] The method may include extracting spatial clues related to the keywords.
[0009] The method may include using the spatial clues related to the keywords as clustering seeds.
[0010] The acoustic cues may include pitch frequency, pitch intensity, one or more pitch frequency harmonics, and the intensities of the one or more pitch frequency harmonics.
[0011] The method may include associating a reliability attribute with each pitch and determining that a speaker associated with the pitch may be silent when the reliability of the pitch drops below a predefined threshold.
[0012] The clustering may include processing the frequency-transformed samples to provide the acoustic cues and the spatial cues; tracking the time-varying state of the speaker using the acoustic cues; partitioning the spatial cues of each frequency component of the frequency-transformed signal into groups; and assigning the acoustic cues associated with the currently active speaker to each group of frequency-transformed signals.
[0013] The assignment may include calculating the cross-correlation between elements of equal-frequency lines of a time-frequency map for each group of frequency-transformed signals and elements belonging to other lines of the time-frequency map and associated with the group of frequency-transformed signals.
[0014] The tracking may include applying an Extended Kalman Filter.
[0015] The tracking may include applying Multiple Hypothesis Tracking.
[0016] The tracking may include applying a Particle Filter.
[0017] The partitioning may include assigning a single frequency component associated with a single time frame to a single speaker.
[0018] The method may include monitoring at least one monitored acoustic feature of speech speed, speech intensity, and emotional utterance.
[0019] The method may include feeding the at least one monitored acoustic feature into an Extended Kalman Filter.
[0020] The frequency-transformed samples may be arranged in a plurality of vectors, one vector for each microphone in the microphone array; wherein the method may include calculating an intermediate vector by weighted averaging of the plurality of vectors; and searching for acoustic cue candidates by ignoring elements of the intermediate vector having values that may be below a predefined threshold.
[0021] The method may include determining the predefined threshold as three times the standard deviation of the noise.
[0022] A non-transitory computer-readable medium can be provided that stores instructions which, when executed by a computerized system, cause the computerized system to: receive or generate a sound sample that represents a sound signal received by a microphone array during a given time period; perform a frequency transformation on the sound sample to provide a frequency-transformed sample; cluster the frequency-transformed sample according to speakers to provide speaker-related clusters, where the clustering can be based on (i) spatial cues related to the received sound signal and (ii) acoustic cues related to the speakers; determine a relative transfer function for each of the speakers to provide speaker-related relative transfer functions; apply a multi-input multi-output (MIMO) beamforming operation on the speaker-related relative transfer functions to provide a beamformed signal; and perform an inverse frequency transformation on the beamformed signal to provide a speech signal.
[0023] The non-transitory computer-readable medium can store instructions for generating the acoustic cues related to the speakers.
[0024] The generation of the acoustic cues can include searching for keywords in the sound sample; and
[0025] extracting the acoustic cues from the keywords.
[0026] The generation of the acoustic cues can include searching for keywords in the sound sample; and
[0027] extracting the acoustic cues from the keywords.
[0028] The non-transitory computer-readable medium can store instructions for extracting spatial cues related to the keywords.
[0029] The non-transitory computer-readable medium can store instructions for using the spatial cues related to the keywords as clustering seeds.
[0030] The acoustic cues can include pitch frequency, pitch intensity, one or more pitch frequency harmonics, and the intensity of the one or more pitch frequency harmonics.
[0031] The non-transitory computer-readable medium can store instructions for associating a reliability attribute with each pitch and determining that a speaker associated with the pitch can be silent when the reliability of the pitch drops below a predefined threshold.
[0032] The clustering may include processing the frequency-transformed samples to provide the acoustic cues and the spatial cues; tracking a time-varying state of a speaker using the acoustic cues; partitioning the spatial cues of each frequency component of the frequency-transformed signal into groups; and assigning the acoustic cues associated with a currently active speaker to each group of frequency-transformed signals.
[0033] The assignment may include calculating a cross-correlation between elements of an isofrequency row of a time-frequency map and elements belonging to other rows of the time-frequency map and that may be associated with the group of frequency-transformed signals for each group of frequency-transformed signals.
[0034] The tracking may include applying an extended Kalman filter.
[0035] The tracking may include applying multiple hypothesis tracking.
[0036] The tracking may include applying a particle filter.
[0037] The partitioning may include assigning a single frequency component associated with a single time frame to a single speaker.
[0038] The non-transitory computer-readable medium may store instructions for monitoring acoustic features for at least one of monitoring speech speed, speech intensity, and emotional expression.
[0039] The non-transitory computer-readable medium may store instructions for feeding the at least one monitored acoustic feature into an extended Kalman filter.
[0040] The frequency-transformed samples may be arranged in a plurality of vectors, one vector for each microphone in the microphone array; wherein the non-transitory computer-readable medium may store instructions for calculating an intermediate vector by weighted averaging of the plurality of vectors and for searching for acoustic cue candidates by ignoring elements of the intermediate vector having values that may be below a predefined threshold.
[0041] The non-transitory computer-readable medium may store instructions for determining the predefined threshold as three times the standard deviation of noise.
[0042] A computerized system can be provided, which can include a microphone array, a memory unit, and a processor. The processor can be configured to receive or generate sound samples that represent sound signals received by the microphone array during a given time period; perform a frequency transformation on the sound samples to provide frequency-transformed samples; cluster the frequency-transformed samples according to speakers to provide speaker-related clusters, where the clustering can be based on (i) spatial cues related to the received sound signals and (ii) acoustic cues related to the speakers; determine a relative transfer function for each of the speakers to provide speaker-related relative transfer functions; apply a multi-input multi-output (MIMO) beamforming operation on the speaker-related relative transfer functions to provide beamformed signals; perform an inverse frequency transformation on the beamformed signals to provide speech signals; and where the memory unit can be configured to store at least one of the sound samples and the speech signals.
[0043] The computerized system may not include the microphone array, but may receive a signal representing the sound signals received by the microphone array during a given time period from the microphone array.
[0044] The processor can be configured to generate the acoustic cues related to the speakers.
[0045] The generation of the acoustic cues can include searching for keywords in the sound samples; and
[0046] extracting the acoustic cues from the keywords.
[0047] The processor can be configured to extract spatial cues related to the keywords.
[0048] The processor can be configured to use the spatial cues related to the keywords as clustering seeds.
[0049] The acoustic cues can include pitch frequency, pitch intensity, one or more pitch frequency harmonics, and the intensity of the one or more pitch frequency harmonics.
[0050] The processor can be configured to associate a reliability attribute with each pitch and determine that the speaker associated with the pitch can be silent when the reliability of the pitch drops below a predefined threshold.
[0051] The processor may be configured to perform clustering by processing the frequency-transformed samples to provide the acoustic cues and the spatial cues; track the time-varying state of a speaker using the acoustic cues; segment the spatial cues of each frequency component in the frequency-transformed signal into groups; and assign the acoustic cues associated with the currently active speaker to each group of frequency-transformed signals.
[0052] The processor may be configured to assign by calculating the cross-correlation between elements of an isofrequency row of a time-frequency map and elements belonging to other rows of the time-frequency map and that may be associated with the group of frequency-transformed signals for each group of frequency-transformed signals.
[0053] The processor may be configured to track by applying an extended Kalman filter.
[0054] The processor may be configured to track by applying multiple hypothesis tracking.
[0055] The processor may be configured to track by applying a particle filter.
[0056] The processor may be configured to segment by assigning a single frequency component associated with a single time frame to a single speaker.
[0057] The processor may be configured to monitor at least one monitored acoustic feature of speech speed, speech intensity, and emotional expression.
[0058] The processor may be configured to feed the at least one monitored acoustic feature into an extended Kalman filter.
[0059] The frequency-transformed samples may be arranged in a plurality of vectors, one vector for each microphone in the microphone array; wherein the processor may be configured to calculate an intermediate vector by performing a weighted average of the plurality of vectors; and search for acoustic cue candidates by ignoring elements of the intermediate vector having values that may be below a predefined threshold.
[0060] The processor may be configured to determine the predefined threshold as three times the standard deviation of the noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] To understand the present invention and to see how it may be put into practice, a preferred embodiment will now be described, by way of non-limiting example only, with reference to the accompanying drawings.
[0062] Figure 1 Illustrates multipath;
[0063] Figure 2 Illustrates an example of a method;
[0064] Figure 3 illustrates Figure 2 an example of the clustering step of the method of;
[0065] Figure 4 illustrates an example of pitch detection on a time-frequency diagram;
[0066] Figure 5 illustrates an example of a time-frequency-Cue map;
[0067] Figure 6 illustrates an example of a voice recognition chain in offline training;
[0068] Figure 7 illustrates an example of a voice recognition chain in real-time training;
[0069] Figure 8 illustrates an example of a training mechanism; and
[0070] Figure 9 illustrates an example of a method. DETAILED DESCRIPTION
[0071] Any reference to a system, with appropriate modifications, shall apply to the method performed by the system and / or a non-transitory computer-readable medium storing instructions that, when executed by the system, cause the system to perform the method.
[0072] Any reference to a method, with appropriate modifications, shall apply to a system configured to perform the method and / or a non-transitory computer-readable medium storing instructions that, when executed by the system, cause the system to perform the method.
[0073] Any reference to a non-transitory computer-readable medium, with appropriate modifications, shall apply to the method performed by a system and / or a system configured to execute instructions stored in the non-transitory computer-readable medium.
[0074] The term "and / or" is additional or alternative.
[0075] The term "system" refers to a computerized system.
[0076] Speech enhancement methods focus on extracting the speech signal from a desired source (speaker) when the signal is corrupted by noise and other speakers. In a free-field environment, spatial filtering in the form of directional beamforming is effective. However, in a reverberant environment, the speech from each source trails (smears) in several directions, not necessarily continuously, thus reducing the advantages of a conventional beamformer. Using a transfer function (TF)-based beamformer to solve this problem, or using the relative transfer function (RTF) as the TF itself is a promising direction. However, in a multi-speaker environment, when speech signals are captured simultaneously, the ability to estimate the RTF for each speaker remains a challenge. A solution is provided that involves tracking acoustic and spatial cues to cluster co-occurring speakers, thus facilitating the estimation of the speaker's RTF in a reverberant environment.
[0077] A clustering algorithm for speakers is provided that assigns each frequency component to its original speaker, especially in a multi-speaker reverberant environment. This provides the necessary condition for the RTF estimator to work properly in a multi-speaker reverberant environment. The estimation of the RTF matrix is then used to calculate the weighted vector of a transfer function-based linearly constrained minimum variance (TF-LCMV) beamformer (see Equation (10) below), and thus the necessary conditions for TF-LCMV operation are met. It is assumed that each human speaker is assigned a different pitch, such that pitch is a bijective indicator for the speaker. Multi-pitch detection is known to be a challenging task, especially in a noisy, reverberant multi-speaker environment. To address this challenge, the W-Disjoint Orthogonality (W-DO) assumption is adopted, and a set of spatial cues (such as signal strength, azimuth, and elevation) are used as additional features. An Extended Kalman Filter (EKF) is used to track the acoustic cue - pitch value over time to overcome temporarily inactive speakers and pitch variations, and spatial cues are used to segment the last L frequency components and assign each frequency component to a different source. The results of the EKF and the segmentation are combined by cross-correlation in order to cluster the frequency components according to a specific speaker with a specific pitch.
[0078] Figure 1Describes the path along which the frequency components of a voice signal propagate from a human speaker 11 to a microphone array 12 in a reverberant environment. The wall 13 and other elements 14 in the environment reflect the impinging signal with attenuation and reflection angles depending on the material and texture of the wall. Different frequency components of human speech may take different paths. They may be the direct path 15 on the shortest path between the human speaker 11 and the microphone array 12, or indirect paths 16, 17. Note that the frequency components may propagate along one or more paths.
[0079] Figure 2 Describes an algorithm. The signal is acquired by a microphone array 201, which includes M ≥ 2 microphones, where M = 7 microphones is an example. The microphones can be deployed in a series of constellations such as equidistantly on a straight line, on a circle, or on a sphere, or even non-equidistantly to form an arbitrary shape. The signals from each microphone are sampled, digitized, and stored in M frames, each frame containing T consecutive samples 202. The size of the frame T can be chosen to be large enough so that the short-time Fourier transform (STFT) is accurate, but short enough so that the signal is stationary along the equivalent duration. For a sampling rate of 16 kHz, a typical value of T is 4096 samples, in other words, the frame corresponds to 1 / 4 second. Usually, consecutive frames overlap with each other to improve the tracking after the characteristics of the signal change over time. A typical overlap is 75%, in other words, a new frame is initiated every 1024 samples. The range of T can be, for example, between 0.1 second - 2 seconds - thus sampling 1024 - 32768 for a sampling rate of 16 kHz. The said samples are also referred to as sound samples representing the sound signal received by the microphone array during the time period T.
[0080] By applying the Fourier transform or a variant of the Fourier transform such as the short-time Fourier transform (STFT), the constant-Q transform (CQT), the logarithmic Fourier transform (LFT), a filter bank, etc., each frame is transformed to the frequency domain in 203. Multiple techniques - such as windowing and zero-padding - can be applied to control the framing effect. The result of 203 is M complex-valued vectors of length K. For example, if the array includes 7 microphones, 7 vectors are prepared, and the vectors are indexed with the frame time registered. K is the number of frequency bins (frequency windows) and is determined by the frequency transform. For example, when using the ordinary STFT, K = T which is the length of the buffer. The output of step 203 can be referred to as the frequency-transformed signal.
[0081] In 204, the voice signals are clustered according to different speakers. The clustering can be referred to as speaker-related clustering. Different from the prior art works that cluster speakers only based on direction, 204 processes multiple speakers in a reverberant room such that signals from different directions can be assigned to the same speaker due to direct paths and indirect paths. The proposed solution suggests using a set of acoustic cues (such as pitch frequency and intensity and their harmonic frequencies and intensities) on top of a set of spatial cues (such as the direction (azimuth and elevation) and intensity of the signal in one of the microphones). One or more of the pitch and spatial cues are used as the state vector for a tracking algorithm (such as a Kalman filter and its variants, multiple hypothesis tracking (MHT), or particle filter), and the tracking algorithm is used to track this state vector and assign each tracking to a different speaker.
[0082] All these tracking algorithms use a model that describes the dynamics of the state vector over time such that when the measurement results of the state vector are lost or corrupted by noise, the tracking algorithm uses this dynamic model to compensate for it and at the same time updates the model parameters. The output of this stage is a vector that assigns each frequency component at a given time Figure 3 to each speaker. 204 is further elaborated in
[0083] In 205, an RTF estimator is applied to the data in the frequency domain. The result of this stage is a set of RTFs, and each RTF is registered to an associated speaker. The registration process is completed using the clustering array from the speaker clustering 204. The set of RTFs is also referred to as speaker-related relative transfer functions.
[0084] The MIMO beamformer 206 reduces the energy of noise and the energy of interfering signals relative to the energy of the desired voice signal through spatial filtering. The output of step 206 can be referred to as the beamformed signal. Then the beamformed signal is forwarded to the inverse frequency transform 207 to create a continuous voice signal in the form of a sample stream, which is then passed to other components such as speech recognition, communication systems, and recording devices 208.
[0085] In a preferred embodiment of the present invention, keyword spotting 209 can be used to improve the performance of the clustering block 204. The frames from 202 are searched for predefined keywords (e.g., "hello Alexa" or "ok Google"). Once a keyword is identified in the stream of frames, the acoustic cues of the speaker are extracted, such as pitch frequency and intensity and their harmonic frequencies and intensities. In addition, the characteristics of the path by which each frequency component reaches the microphone array 201 are extracted. These characteristics are used as seeds for the clustering of the speaker clustering 204 for the desired speaker. A seed is an initial guess about the initial parameters of the clustering. For example, for the centroid, radius, and statistical values of a centroid-based clustering algorithm such as K-means, PSO, and 2KPM. Another example is the base of the subspace for subspace-based clustering.
[0086] Figure 3 A speaker clustering algorithm is described. It is assumed that each speaker is assigned a different set of acoustic cues, e.g., pitch frequency and intensity and their harmonic frequencies and intensities, such that the set of acoustic cues is a bijective indicator for the speaker. Acoustic cue detection is known to be a challenging task, especially in noisy, reverberant multi-speaker environments. To address this challenge, spatial cues in the form of, for example, signal strength, azimuth, and elevation are used. Filters such as particle filters and extended Kalman filters (EKF) are used to track the acoustic cues over time to overcome temporarily inactive speakers and changes in the acoustic cues, and the spatial cues are used to split the frequency components between different sources. The results of the EKF and the splitting are combined by cross-correlation in order to cluster the frequency components according to a specific speaker with a specific pitch.
[0087] As an example of a preferred embodiment, in 31, potential acoustic cues in the form of pitch frequencies are detected. First, a time-frequency map is prepared using a frequency transform of the buffers from each microphone, which is calculated in 203. Next, the absolute value of each of the M complex-valued vectors of length K is weighted and averaged with some weighting factors, which can be determined to reduce artifacts in some microphones. The result is a single real vector of length K. In this vector, values above a given threshold μ are extracted, while the remaining elements are discarded. The threshold μ is typically adaptively selected to be three times the standard deviation of the noise, but not less than a constant value that depends on the electrical parameters of the system and, in particular, on the number of significant bits of the sampled signal. Values with frequency indices in the range [k_min, k_max] are defined as candidates for pitch frequencies. The variables k_min and k_max are typically 85 Hz and 2550 Hz, respectively, since a typical adult male will have a fundamental frequency ranging from 85 Hz to 1800 Hz, and a typical adult female has a fundamental frequency ranging from 165 Hz to 2550 Hz. Then each pitch candidate is verified by searching for its higher harmonics. The presence of the second and third harmonics may be a prerequisite for the candidate pitch to be detected as a legitimate pitch with a reliability R (e.g., R = 10). If higher harmonics are present (e.g., the fourth and fifth), the reliability of the pitch can be increased - for example, the reliability can be doubled for each harmonic. In Figure 4 an example can be found. In a preferred embodiment of the present invention, the pitch 32 of the desired speaker is supplied by 210 using a keyword uttered by the desired speaker. The supplied pitch 32 is added to the list with the highest possible reliability (e.g., R = 1000).
[0088] In 33, an Extended Kalman Filter (EKF) is applied to the pitch from 31. As pointed out in the Wikipedia entry on the Extended Kalman Filter (www.wikipedia.org / wiki / extended_Kalman_filter), the Kalman filter has a state transition equation and an observation model. For discrete calculations, the state transition equation is:
[0089] x k = f(x k-1 , u k ) + w k (1)
[0090] And for discrete calculations, the observation model is:
[0091] zk = h(x k ) + v k (2)
[0092] where x k is the state vector, which contains (partially) the parameters that describe the state of the system, u k is the vector of external inputs that provide information about the state of the system, w k and v k are the process noise and the observation noise. The time updater of the extended Kalman filter can use the prediction equation to predict the next state, and the detected pitch can be updated by comparing the actual measurement with the predicted measurement using an equation of the following type:
[0093] y k = z k - h(x k|k+1 ) (3)
[0094] where z k is the detected pitch, and y k is the error between the measurement and the predicted pitch.
[0095] In 33, each trajectory can start from a detected pitch, followed by the model f(x k, u k ), which reflects the temporal behavior of the pitch, which may become higher or lower due to emotion. The input to this model can be the past state vector x k (one or more), and any external input u k that affects the dynamics of the pitch, such as the speech rate, speech intensity, and emotional expression. The elements of the state vector x can quantitatively describe the pitch. For example, the state vector of the pitch may particularly include the pitch frequency, the intensity of the first harmonic, and the frequencies and intensities of the higher harmonics. The vector function f(x k, u k ) can be used to predict the state vector x at a given time k + 1 before the current time. An exemplary implementation of the dynamic model in the EKF can include the time update equation (also known as the prediction equation), as described in the book "Lessons in Digital Estimation Theory" by Jerry M. Mendel, which is incorporated herein by reference.
[0096] For example, considering the 3 - tuple state vector:
[0097]
[0098] where f kis the frequency of the pitch (1st harmonic) at time k, a k is the intensity of the pitch (1st harmonic) at time k, and b k is the intensity of the 2nd harmonic at time k.
[0099] An exemplary state vector model for pitch can be:
[0100]
[0101] which describes a model that assumes a constant pitch at all times. In a preferred embodiment of the present invention, the speech speed, speech intensity, and emotional expression using speech recognition algorithms known in the art are continuously monitored to provide an external input u to improve the time update phase of the EKF k . The method of emotional expression is known in the art. For example, see "New Features for Emotional Speech Recognition" by Palo et al.
[0102] Each track is given a reliability field that is inversely proportional to the time for which the track evolves using only time updates. When the reliability of a track drops below a certain reliability threshold ρ, such as representing an undetected pitch for 10 seconds, the track is defined as dead (no longer relevant), meaning the corresponding speaker is not active. On the other hand, when a new measurement result (pitch detection) that cannot be assigned to any existing track appears, a new track is initiated.
[0103] In 34, spatial cues are extracted from M frequency-transformed frames. As in 31, the most recent L vectors are saved for analysis using temporal correlation. The result is a time-frequency-cue (TFC) map, which is a 3D array of size LxKxP (where P = M - 1) for each of the M microphones. Figure 5 Describes the TFC.
[0104] In 35, the spatial cues of each frequency component in the TFC are segmented. The idea is that along the L frames, a frequency component may originate from different speakers, and this can be observed by comparing the spatial cues. However, due to the W-DO assumption, it is assumed that at a single frame time l, the frequency component originates from a single speaker. Any known method for clustering in the literature (such as K-Nearest Neighbor (KNN)) can be used to perform the segmentation. Clustering assigns an index to each cell in A which indicates to which cluster the cell (k, l) belongs.
[0105] In 36, the frequency components of the signal are grouped such that each frequency component is assigned to a particular pitch in the list of pitches tracked by the EKF and is active due to its reliability. This is done by computing the sample cross-correlation between the k-th row of the time-frequency diagram (see Figure 4 ) assigned to one of the pitches and all values of (j, l) in the other rows of this time-frequency diagram having a particular clustering index c 0 . This is done for each clustering index. The sample cross-correlation is given by:
[0106]
[0107] where A is the time-frequency diagram, k is the index of the row belonging to one of the pitches, j is any other row of A, and L is the number of columns of A. After computing the cross-correlation between each pitch and each of the clusters in the other rows, the cluster c 1 in row j 1 with the highest cross-correlation is grouped with the corresponding pitch, and then the cluster c 2 in row j 2 with the second highest cross-correlation is grouped with the corresponding pitch, and so on. This process is repeated until the sample cross-correlation drops below a certain threshold κ, which can be adaptively set to, for example, 0.5x (the average energy of the signal at a single frequency). The result of 35 is a set of frequency groups to which the corresponding pitch frequencies are assigned.
[0108] Figure 4 Describes an example of pitch detection on a time-frequency diagram. 41 is the time axis, which is represented by the parameter , and 42 is the frequency axis, which is described by the parameter k. Each column in this two-dimensional array is a K-length real-valued vector extracted in 31 after averaging the absolute values of the M frequency-transformed buffers at time . For performing correlation analysis in time, the L nearest vectors are saved in a two-dimensional array of size KxL. In 43, two pitches are represented by diagonals in different directions. Due to the presence of 4th harmonics, the pitch k = 2 - whose harmonics are at k = 4, 6, 8 - has a reliability R = 20, and the pitch at k = 3 - whose harmonics are at k = 6, 9 - has a reliability R = 10. In 44, the pitch at k = 3 is inactive and only k = 2 is active. However, since the 4th harmonics are not detected (below the threshold μ), the reliability of the pitch at k = 2 is reduced to R = 10. In 45, the pitch at k = 3 is active again and k = 2 is inactive. In 46, a new pitch candidate at k = 4 appears, but only its 2nd harmonic is detected. Therefore, it is not detected as a pitch. In 47, the pitch at k = 3 is inactive and no pitch is detected.
[0109] Figure 5 describes a TFC graph, whose axes are frame index (time) 51, frequency components 52, and spatial cues 53, which may be, for example, complex values expressing the direction (azimuth and elevation) in which each frequency component arrives and the component intensity. When a frame with index is processed and passed into the frequency domain, for each frequency element a vector of M complex numbers is received. Up to M - 1 spatial cues are extracted from each vector. In the example of the direction and intensity of each frequency component, this may be done using any direction-finding algorithm known in the art for array processing (such as MUSIC or ESPRIT). The result of this algorithm is a set of up to M - 1 directions in three-dimensional space, each expressed by two angles and the estimated intensity of the arriving signal, the intensity p = 1,..., P ≤ M - 1. The cues are arranged in the TFC graph such that at the cell indexed by .
[0110] Appendix
[0111] The performance of the speech enhancement module depends on the ability to filter out all interfering signals and leave only the desired speech signal. Interfering signals may be, for example, other speakers, noise from air conditioners, music, motor noise (e.g., in a car or an airplane), and large crowd noise also known as "cocktail party noise". The performance of the speech enhancement module is usually measured by its ability to improve the speech-to-noise ratio (SNR) or speech-to-interference ratio (SIR), which respectively reflect the ratio of the power of the desired speech signal to the total power of the noise and the ratio of the power of the desired speech signal to the total power of other interfering signals (usually in dB).
[0112] When the acquisition module contains a single microphone, the method is called single-microphone speech enhancement and is usually based on the statistical characteristics of the signal itself in the time-frequency domain, such as single channel spectral subtraction, spectral estimation using minimum variance distortionless response (MVDR), and echo cancellation. When more than a single microphone is used, the acquisition module is usually called a microphone array, and the method is called multi-microphone speech enhancement. Many of these methods utilize the differences between the signals captured simultaneously by the microphones. One well-established method is beamforming, which calculates the total of the signals from the microphones after multiplying each signal by a weighting factor. The goal of the weighting factor is to average out the interfering signals in order to condition the signal of interest.
[0113] In other words, beamforming is a method of creating a spatial filter that algorithmically increases the power of signals transmitted from a given location in space (desired signals from a desired speaker) and decreases the power of signals transmitted from other locations in space (interference signals from other sources), thereby increasing the SIR at the output of the beamformer.
[0114] The delay-and-sum beamformer (DSB) involves using the weighting factors of the DSB, which consist of the counter delays implied by the different paths along which the desired signal propagates from its source to each microphone in the array. The DSB is restricted to signals each from a single direction, such as in a free-field environment. Thus, in a reverberant environment, where signals from the same source propagate along different paths to the microphones and arrive at the microphones from multiple directions, the DSB performance is usually insufficient.
[0115] To mitigate the drawbacks of the DSB in a reverberant environment, the beamformer can use a more complex acoustic transfer function (ATF), which represents the direction (azimuth and elevation) of each frequency component from a given source to a specific microphone. The single arrival direction (DOA) assumed by the DSB and other DOA-based methods is generally not applicable in a reverberant environment, where components of the same speech signal arrive from different directions. This is due to the different frequency responses of the physical elements (such as walls, furniture, and people) in a reverberant environment. The ATF in the frequency domain is a vector that assigns a complex number to each frequency in the Nyquist bandwidth. The absolute value represents the path gain associated with this frequency, and the phase indicates the phase added to the frequency component along this path.
[0116] Estimating the ATF between a given point in space and a given microphone can be done by using a speaker located at that given point and transmitting a known signal. By simultaneously acquiring the signal from the speaker as input and the signal from the output of the microphone, one can easily estimate the ATF. The speaker can be located at one or more positions where a human speaker may reside during the operation of the system. This method creates a map of the ATF for each point in space or more practically for each point on a grid. The ATF for points not included in the grid is approximated using interpolation. Nevertheless, this method has serious drawbacks. First, the system needs to be calibrated for each installation, making this method impractical. Second, the acoustic differences between a human speaker and an electronic speaker, which cause the measured ATF to deviate from the actual ATF. Third, the complexity of measuring a large number of ATFs, especially when the direction of the speaker is also considered, and fourth, the possible errors due to environmental changes.
[0117] A more practical alternative to the ATF is the relative transfer function (RTF) as a remedy for the drawbacks of the ATF estimation method in practical applications. The RTF is the difference between the ATFs between a given source and two of the microphones in the array, which in the frequency domain takes the form of the ratio between the spectral representations of the two ATFs. Like the ATF, the RTF in the frequency domain assigns complex numbers to each frequency. The absolute value is the gain difference between the two microphones, which is usually close to unity when the microphones are close to each other, and under some conditions, the phase reflects the angle of incidence of the source.
[0118] The transfer function-based linearly constrained minimum variance (TF-LCMV) beamformer can reduce noise while limiting speech distortion by minimizing the output energy under the constraint that the speech component in the output signal equals the speech component in one of the microphone signals in multi-microphone applications. If N = N d + N i sources, considering the problem of extracting N i desired speech sources contaminated by N d interfering sources and stationary noise. Each of the signals involved propagates through the acoustic medium and is then picked up by an arbitrary array consisting of M microphones. The signal of each microphone is segmented into frames of length T, and the FFT is applied to each frame. In the frequency domain, let us denote the k-th frequency component of the m-th microphone and the n-th source of the and th frame respectively. Similarly, the ATF between the n-th source and the m-th microphone is and the noise at the m-th microphone is The received signal in matrix form is given by:
[0119]
[0120] where is the sensor vector, is the source vector, is the ATF matrix such that and is the additive stationary noise uncorrelated with any source. Equivalently, (7) can be formulated using the RTF. Without loss of generality, the RTF of the n-th speech source can be defined as the ratio between the n-th speech component at the m-th microphone and its corresponding component at the first microphone, i.e., The signal in (7) can be formulated using the RTF matrix such that In vector notation:
[0121]
[0122] wherein is the source signal of the change.
[0123] Considering the array measurement results it is necessary to estimate the mixture of N d desired sources. By applying a beamformer to the microphone signals the extraction of the desired signals can be achieved. Assume M≥N can be selected to meet the LCMV criterion:
[0124]
[0125] wherein is the power spectral density (PSD) matrix of, and is the constraint vector.
[0126] A possible solution to (9) is:
[0127]
[0128] Based on (7) and (8) and the set constraints, the component of the desired signal at the beamformer output is given by In other words, the output of the beamformer is a mixture of the components of the desired signal measured by the first (reference) microphone.
[0129] From the th set of the RTF and for each frequency component k, up to a set of M - 1 sources - with the incident angle p = 1,.., P≤M - 1 and the elevation angle - can be extracted using, for example, an algorithm based on the phase difference together with the intensity obtained from the microphone defined as the reference microphone among the microphones. These triples are generally referred to as spatial cues.
[0130] TF - LCMV is a suitable method for extracting M - 1 speech sources impinging on an array of M sensors from different positions including in a reverberant environment. However, a necessary condition for TF - LCMV to work is that the RTF matrix whose columns are the RTF vectors of all active sources in the environment is known and available for TF - LCMV. This requires the association of each frequency component with its source speaker.
[0131] Multiple methods can be used to assign sources to signals without complementary information. The main family of methods is known as blind source separation (BSS), which recovers unknown signals or sources from the observed mixtures. A key weakness of BSS in the frequency domain is that, at each frequency, the column vectors of the mixing matrix (estimated by BSS) are randomly permuted, and without knowledge of this random permutation, it becomes difficult to combine the results across frequencies, as disclosed.
[0132] BSS can be assisted by pitch information. However, the gender of the speaker must be a-priory.
[0133] BSS can be used in the frequency domain, while using the maximum-magnitude method to resolve the ambiguity of the estimated mixing matrix, which assigns a particular column of the mixing matrix to the source corresponding to the largest element in the vector. Nevertheless - this method depends largely on the spectral distribution of the sources, as it is assumed that the strongest component at each frequency indeed belongs to the strongest source. However, this condition is not often met, as different speakers may introduce intensity peaks at different frequencies. Alternatively, source activity detection, also known as voice activity detection (VAD), can be used such that information about the sources active at a particular time is used to resolve the ambiguity in the mixing matrix. The drawback of VAD is that it is not able to robustly detect voice pauses, especially in a multi-speaker environment. In addition, this method is only effective when no more than a single speaker joins the conversation at a time, requires a relatively long training period, and is sensitive to movement during this period.
[0134] The TF-LCMV beamformer and its extended forms can be used together with a binaural cue generator for a binaural speech enhancement system. Acoustic cues are used to separate the speech components from the noise components in the input signal. This technique is based on the auditory scene analysis theory 1 , which suggests using unique perceptual cues to cluster signals from different speech sources in a "cocktail party" environment. Examples of the primitive grouping cues that can be used for speech separation include common onsets / offsets across frequency bands, pitch (fundamental frequency), the same location in space, temporal and spectral modulation, pitch and energy continuity and smoothness. However, an implicit assumption of this method is that all components of the desired speech signal have almost the same direction. In other words, an almost free-field condition, saving the influence of the head-shadow effect, which is proposed to be compensated by using head-related transfer functions. This is less likely to occur in a reverberant environment.
[0135] Note that even when multiple speakers are active simultaneously, the spectral content of the speakers does not overlap at most time - frequency points. This is known as W - separated orthogonality, or simply W - DO for short. This can be demonstrated by the sparsity of the speech signal in the time - frequency domain. According to this sparsity, the probability of simultaneous activity of two speakers at a particular time - frequency point is very low. In other words, in the case of multiple simultaneous speakers, each time - frequency point is most likely to correspond to the spectral content of one of the speakers.
[0136] W - DO can be used to facilitate BSS by defining a particular class of signals that are to some extent W - DO. This can be used only when first - order statistics are needed, which is computationally economical. Moreover, if the sources are W - DO and do not occupy the same spatial position, it is possible to de - mix any number of signal sources using only two microphones. However, this method assumes that the underlying mixing matrix is the same across all frequencies. This assumption is necessary for using the histogram of the estimated mixing coefficients across different frequencies. However, this assumption generally does not hold in a reverberant environment and holds only in a free field. The extension of this method to the multipath case is limited to negligible energy from the multipath or to a sufficiently smooth convolutional mixing filter such that the histogram is tailed but maintains a single peak. This assumption also does not hold in a reverberant environment where the differences between different paths are usually too large to create a smooth histogram.
[0137] It has been found that the proposed solution performs in a reverberant environment and does not have to rely on unnecessary assumptions and constraints. The solution can operate even without prior information, even without a large training process, even without constraining the estimation of the attenuation and delay of a given source at each frequency to a single point in the attenuation - delay space, even without constraining the estimated values of the attenuation - delay values of a single source to create a single cluster, and even without limiting the number of mixed voices to two.
[0138] Source separation into a speech recognition engine
[0139] A voice user interface (VUI) is an interface between a human speaker and a machine. The VUI uses one or more microphones to receive speech signals and converts the speech signals into digital signatures, typically by transcribing the speech signals into text, which is used to infer the speaker's intent. Then, the machine can respond to the speaker's intent based on the application the machine is designed for.
[0140] A key component of a VUI is an automatic speech recognition engine (ASR) that converts digitized speech signals into text. The performance of the ASR, in other words, the accuracy with which the text describes the acoustic speech signal, depends to a large extent on the match between the input signal and the requirements of the ASR. Therefore, other components of the VUI are designed to enhance the acquired speech signal before feeding it to the ASR. Such components may be noise suppression, echo cancellation, and source separation, to name just a few.
[0141] One of the key components in speech enhancement is source separation (SS), which is intended to separate speech signals from several sources. Given an array of two or more microphones, the signal acquired by each of the microphones is a mixture of all the speech signals in the environment plus other interferences (such as noise and music). The SS algorithm takes the mixed signals from all the microphones and decomposes them into their components. In other words, the output of source separation is a set of signals, each representing the signal of a particular source, whether it is the speech signal from a particular speaker or music or even noise.
[0142] There is an increasing need to improve source separation.
[0143] Figure 6Illustrates an example of a speech recognition chain in offline training. The chain typically includes a microphone array 511, which provides a set of digitized acoustic signals. The number of digitized acoustic signals is equal to the number of microphones that make up the array 511. Each digitized acoustic signal contains a mixture of all acoustic sources near the microphone array 511, whether human speakers, synthetic speakers such as TV, music, and noise. The digitized acoustic signals are delivered to a preprocessing stage 512. The purpose of the preprocessing stage 512 is to improve the quality of the digitized acoustic signals by removing interferences such as echo, reverberation, and noise. The preprocessing stage 512 is typically performed using a multi-channel algorithm that employs statistical correlations between the digitized acoustic signals. The output of the preprocessing stage 512 is a set of processed signals, typically having the same number of signals as the digitized acoustic signals at the input to this stage. The set of processed signals is forwarded to a source separation (SS) stage 513, which aims to extract acoustic signals from each source near the microphone array. In other words, the SS stage 513 takes a set of signals, where each signal is a different mixture of acoustic signals received from different sources, and creates a set of signals such that each signal mainly contains a single acoustic signal from a single specific source. Source separation of the speech signals can be performed using geometric considerations of the deployment of sources such as beamforming or by considering characteristics of the speech signals such as independent component analysis. The number of separated signals is typically equal to the number of active sources near the microphone array 511, but less than the number of microphones. The set of separated signals is forwarded to a source selector 514. The purpose of the source selector is to select the relevant speech source whose speech signal should be recognized. The source selector 514 can use a trigger word detector such that the source that emits a predefined trigger word sound is selected. Alternatively, the source selector 514 can consider the position of the sources near the microphone array 511, such as a predefined direction relative to the microphone array 511. In addition, the source selector 514 can use a predefined acoustic signature of the speech signal to select the source that matches this signature. The output of the source selector 514 is a single speech signal, which is forwarded to a speech recognition engine 515. The speech recognition engine 515 converts the digitized speech signal into text. There are many methods known in the art for speech recognition, most of which are based on extracting features from the speech signal and comparing these features with a predefined vocabulary. The main output of the speech recognition engine 515 is a text string 516 associated with the input speech signal. In offline training, a predefined text 518 sound is emitted to the microphone. An error 519 is calculated by comparing the output of the ASR 516 with this text. The comparison 517 can be performed using simple word counting or more complex comparison methods that consider the meaning of the words and appropriately weight the error detection for different words. Then the SS 513 uses the error 519 to modify a set of parameters to find the values that minimize the error.This can be done by any supervised estimation or optimization method, such as least squares, stochastic gradient, neural networks (NN) and their variants.
[0144] Figure 7 Illustrates an example of a speech recognition chain during real-time training, i.e., during the normal operation of the system. During the operation of the VUI, the true text pronounced by the human speaker is unknown, and the supervised error 519 is not available. An alternative is the confidence score 521 developed for real-time applications, which can benefit from knowing the reliability level of the ASR output when there is no reference to the true spoken text. For example, when the confidence score is low, the system may enter an appropriate branch where it has a more direct conversation with the user. There are many methods for estimating the confidence score, and most of them aim for a high correlation with the error that can be calculated when the spoken text is known. During real-time training, the confidence score 521 is converted into a supervised error 519 by an error estimator 522. When the confidence score is highly correlated with the theoretical supervised error, the error estimator 522 can be a simple axis conversion. When the range of the confidence score 521 is from 0 to 100 and the goal is to make it higher, and the range of the supervised error is from 0 to 100 and the goal is to make it lower, a simple axis conversion in the form of estimated error = 100 - confidence score is used as the error estimator 522. The estimated error 519 can be used to train the parameters of the SS, and the same applies to offline training.
[0145] Figure 3 Illustrates the training mechanism of a typical SS 513. The source separator (SS) 513 receives a set of mixed signals from the preprocessing stage 512 and feeds the separated signals to the source selector 514. Generally, source separation of acoustic signals and especially speech signals is done in the frequency domain. The mixed signals from the preprocessing stage 512 are first transformed into the frequency domain 553. This is done by dividing the mixed signals into segments of the same length, with overlapping periods between the resulting segments. For example, when the length of the segment is determined to be 1024 samples and the overlapping period is 25%, each of the mixed signals is divided into segments of 1024 samples each. The set of concurrent segments from different mixed signals is called a batch. Each batch of segments starts 768 samples after the previous batch. Note that the segments across the set of mixed signals are synchronous, i.e., the starting points of all segments belonging to the same batch are the same. The length and overlapping period of the segments in a batch are obtained from the model parameters 552.
[0146] The demixing algorithm 554 separates a batch of segments arriving from the frequency transformation 553. Like many other algorithms, the source separation (SS) algorithm consists of a set of mathematical models - accompanied by a set of model parameters 552. The mathematical models establish the way of operation, such as the way SS copes with physical phenomena (e.g., multipath). The set of model parameters 552 tunes the operation of SS to the specific characteristics of the source signals, to the architecture of the automatic speech recognition engine (ASR) receiving these signals, to the geometry of the environment, and even to the human speaker.
[0147] The demixed batch of segments is forwarded to the inverse frequency transformation 555, in which it is transformed back to the time domain. The same set of model parameters 552 used in the frequency transformation stage 553 is used in the inverse frequency transformation stage 555. For example, overlapping periods are used to reconstruct the output signal in the time domain from the resulting batch. This is done, for example, using the overlap - add method, in which after the inverse frequency transformation, the overlapping intervals are overlapped and added to reconstruct the resulting output signal, possibly using an appropriate weighting function ranging from 0 to 1 across the overlapping regions such that the total energy is conserved. In other words, the overlapping segments from the previous batch fade out while the overlapping segments from the subsequent batch fade in. The output of the inverse frequency transformation block is forwarded to the source selector 514.
[0148] The model parameters 552 are a set of parameters used by the frequency transformation block 553, the demixing block 554, and the inverse frequency transformation block 555. Dividing the mixed signal into segments of the same length - which is done by the frequency transformation 553 - is paced by a clocking mechanism such as a real - time clock. At each pace, each of the frequency transformation block 553, the demixing block 554, and the inverse frequency transformation block 555 extracts parameters from the model parameters 552. These parameters are then substituted into the mathematical models executed in the frequency transformation block 553, the demixing block 554, and the inverse frequency transformation block 555.
[0149] The corrector 551 optimizes the set of model parameters 552 with the aim of reducing the error 519 from the error estimator. The corrector 551 receives the error 519 and the current set of model parameters 552, and outputs the corrected set of model parameters 552. The correction of this set of parameters can be done a priori (offline) or during VUI operation (in real - time). In offline training, the error 519 used to correct the set of model parameters 552 is extracted using a predefined text pronounced by the microphone and comparing the output of the ASR with this text. In real - time training, the error 519 is extracted from the confidence score of the ASR.
[0150] This error is then used to modify the set of parameters to find the value that minimizes the error. This can be done by any supervised estimation or optimization method, preferably a derivative-free method such as golden section search, grid search, and Nelder-Mead.
[0151] The Nelder-Mead method (also known as the downhill simplex method, amoeba method, or polytope method) is a commonly applied numerical method for finding the minimum or maximum value of an objective function in a multi-dimensional space. It is a direct search method (based on function comparison) and is generally applicable to non-linear optimization problems where the derivative may not be known.
[0152] The Nelder-Mead method iteratively finds the local minimum of the error 519 as a function of several parameters. The method starts by determining a set of values for a simplex (a generalized triangle in N dimensions). There is an assumption that a local minimum exists within the simplex. At each iteration, the error at the vertices of the simplex is calculated. The vertex with the largest error is replaced with a new vertex such that the volume of the simplex is reduced. This is repeated until the simplex volume is less than a predefined volume, and the best value is one of the vertices. This process is performed by the corrector 551.
[0153] The golden section search finds the minimum of the error 519 by continuously shrinking the range of values within which the minimum is known to exist. The golden section search requires a strictly unimodal error as a function of the parameters. The operation of shrinking the range is done by the corrector 551.
[0154] The golden section search is a technique used to find the extremum (minimum or maximum) of a strictly unimodal function by continuously shrinking the range of values within which the extremum is known to exist. (www.wikipedia.org).
[0155] The grid search iteratively traverses a set of values associated with one or more of the parameters to be optimized. When more than one parameter is being optimized, each value in the set is a vector of length equal to the number of parameters. For each value, the error 519 is calculated, and the value corresponding to the minimum error is selected. The iteration of traversing the set of values is performed by the corrector 551.
[0156] Grid Search - A traditional way to perform hyperparameter optimization is grid search or parameter scanning, which is an exhaustive search that simply traverses a manually specified subset of the hyperparameter space of a learning algorithm. The grid search algorithm must be guided by some performance metric, typically measured by cross-validation on the training set or evaluation on a held-out validation set. Since the parameter space of a machine learner can include real-valued or unbounded value spaces for some parameters, it may be necessary to manually set bounds and discretize before applying grid search. (www.wikipedia.org).
[0157] All optimization methods need to continuously compute the error 519 on the same set of separated acoustic signals. This is a time-consuming process and may therefore not be continuously performed, but only when the error 519 (which is continuously computed) exceeds a predefined threshold (e.g., 10% error). When this occurs, two methods can be employed.
[0158] One method is to run the optimization in parallel using parallel threads or multi-core in parallel with the normal operation of the system. In other words, there are one or more parallel tasks that execute blocks 513, 514, 515, 522 in parallel with the tasks of the normal operation of the system. In the parallel tasks, a batch of mixed signals with a length of 1 - 2 seconds is obtained from the preprocessing 512 and repeatedly separated 513 and decoded 514, 515 with a set of different model parameters 552. The error 519 is computed for each such loop. In each loop, the corrector 551 selects a set of model parameters according to the optimization method.
[0159] The second method is to run the optimization when there is no speech in the room. A voice activity detection (VAD) algorithm can be used to detect periods of absence of human speech. These periods are used to optimize the model parameters 552 in the same way as in the first method to save the need for parallel threads or multi-core.
[0160] A suitable optimization method should be selected for each parameter in 552. Some methods are applicable to a single parameter, and some are applicable to a group of parameters. Several parameters that affect the performance of speech recognition are suggested below. In addition, optimization methods are suggested based on the characteristics of the parameters.
[0161] Segment length parameter.
[0162] The segment length parameter is related to FFT / IFFT. Generally, ASR using the features of discrete phonemes requires short segments of about 20 milliseconds, while ASR using the features of a series of resulting phonemes uses segments of about 100 - 200 milliseconds. On the other hand, the length of the segment is affected by the scenario, such as the reverberation time of the room. The segment length should be about the reverberation time, which can be up to 200 - 500 milliseconds. Since there is no "sweet point" for the segment length, this value should be optimized for the scenario and ASR. In terms of samples, typical values are 100 - 500 milliseconds. For example, a sampling rate of 8 kHz means a segment length of 800 - 4000 samples. This is a continuous parameter.
[0163] The optimization of this parameter can be done using various optimization methods, such as golden section search or Nelder - Mead together with overlapping periods. When using the golden section search, the input to the algorithm is the minimum and maximum possible lengths, e.g., 10 milliseconds to 500 milliseconds, and the error function 519. The output is the segment length that minimizes the error function 519. When using Nelder - Mead with overlapping periods, the input is a set of three pairs of overlapping period and segment length, e.g., (10 milliseconds, 0%), (500 milliseconds, 10%), and (500 milliseconds, 80%) and the error function 519, and the output is the optimal segment length and the optimal overlapping period.
[0164] Overlapping period
[0165] The overlapping period parameter is related to FFT / IFFT. The overlapping period is used to avoid ignoring phonemes due to segmentation. In other words, phonemes are divided among the resulting segments. Due to the segment length, the overlapping period depends on the features adopted by the ASR. The typical range is 0 - 90% of the segment length. This is a continuous parameter.
[0166] The optimization of this parameter can be done using various optimization methods, such as golden section search or Nelder - mead with the segment length. When using the golden section search, the input to the algorithm is the minimum and maximum possible overlapping periods, e.g., 0% to 90%, and the error function 519. The output is the overlapping period that minimizes the error function 519.
[0167] Window. The window parameters are related to FFT / IFFT. Frequency transformation 553 typically uses windowing to mitigate the effects of segmentation. Some windows such as Kaiser and Chebyshev are parameterized. This means that the effect of the window can be controlled by changing the parameters of the window. The typical range depends on the type of window. This is a continuous parameter. Various optimization methods (such as golden section search) can be used to optimize this parameter. When using the golden section search, the inputs to the algorithm are the minimum and maximum values of the window parameters - which depend on the window type, and the error function 519. For example, for the Kaiser window, the minimum and maximum values are (0, 30). The output is the optimal window parameter.
[0168] Sampling rate
[0169] The sampling rate parameter is related to FFT / IFFT. The sampling rate is one of the key parameters affecting the performance of speech recognition. For example, there are ASRs that exhibit poor results for sampling rates below 16 kHz. Others can work well even at 4 kHz or 8 kHz. Typically, this parameter is optimized once when the ASR is selected. The typical range is 4, 8, 16, 44.1, 48 kHz. This parameter is a discrete parameter. Various optimization methods (such as grid search) can be used to optimize this parameter. The inputs to the algorithm are the values for which the grid search is performed - for example, the sampling rates of (4, 8, 16, 44.1, 48) kHz, and the error function 519. The output is the optimal sampling rate.
[0170] Filtering
[0171] The filtering parameters are related to demixing. Some ASRs use features that represent a limited frequency range. Therefore, filtering the separated signals after source separation 513 can emphasize the specific features used by the ASR, thus improving its performance. Moreover, filtering out the spectral components not used by the ASR can improve the signal-to-noise ratio (SNR) of the separated signals, which in turn can improve the performance of the ASR. The typical range is 4 kHz - 8 kHz. Various optimization methods (such as golden section search) can be used to optimize this parameter. This parameter is continuous. When applying the golden section search, the inputs to the algorithm are the error function 519 and an initial guess for the cut-off frequency part, such as 1000 Hz and 0.5X the sampling rate. The output is the optimal filtering parameter.
[0172] The weighting factor for each microphone. The weighting factor for each microphone is related to de-mixing. In theory, the sensitivities of different microphones on a particular array should be similar, up to 3 dB. However, in practice, the span of the sensitivities of different microphones can be larger. Additionally, the sensitivity of a microphone may change over time due to dust and humidity. The typical range is 0 - 10 dB. This is a continuous parameter. Various optimization methods (such as Nelder - mead) can be used to optimize this parameter whether or not there is a weighting factor for each microphone. When applying the Nelder - mead method, the input to the algorithm is the error function 519 and an initial guess of the simplex vertices. For example, the size of each n - tuple is the number of microphones - N: (1,0,…,0,0), (0,0,…,0,1), and (1 / N,1 / N,...,1 / N). The output is the optimal weighting for each microphone.
[0173] The number of microphones
[0174] The number of microphones is related to de - mixing. The number of microphones affects, on the one hand, the number of sources that can be separated and, on the other hand, the complexity and numerical accuracy. Practical experiments also show that too many microphones may lead to a decrease in the output SNR. The typical range is 4 - 8. This is a discrete parameter. Various optimization methods (such as grid search or Nelder - mead) can be used to optimize this parameter. When applying grid search, the input to the algorithm is the error function 519 and the number of microphones for which the search is performed. For example, 4, 5, 6, 7, 8 microphones. The output is the optimal number of microphones.
[0175] Figure 9 Method 600 is illustrated.
[0176] Method 600 can start from step 610 of receiving or calculating an error related to a speech recognition process applied to a previous output of a source selection process.
[0177] After step 610 can be step 620 of modifying at least one parameter of the source separation process based on the error.
[0178] After step 620 can be step 630 of receiving an audio signal representing audio signals originating from multiple sources and detected by a microphone array.
[0179] After step 630 can be step 640 of performing a source separation process to separate different source audio signals originating from the multiple sources to provide a source - separated signal and sending the source - separated signal to the source selection process.
[0180] After step 640 can be step 630.
[0181] After each or multiple iterations of step 630 and step 640, it can be step 610 (not shown) - where the output of step 640 can be fed into the source selection process and the ASR to provide the previous output of the ASR.
[0182] It should be noted that the initial iterations of step 630 and step 640 can be performed without receiving an error.
[0183] Step 640 can include applying frequency conversion (such as but not limited to FFT), demixing, and applying inverse frequency conversion (such as but not limited to IFFT).
[0184] Step 620 can include at least one of the following:
[0185] a. Modify at least one parameter of the frequency conversion.
[0186] b. Modify at least one parameter of the inverse frequency conversion.
[0187] c. Modify at least one parameter of the demixing.
[0188] d. Modify the length of the segment of the signal representing the audio signal to which the frequency conversion is applied.
[0189] e. Modify the overlap between consecutive segments of the signal representing the audio signal, where the frequency conversion is applied on a segment-by-segment basis.
[0190] f. Modify the sampling rate of the frequency conversion.
[0191] g. Modify the windowing parameter of the window applied by the frequency conversion.
[0192] h. Modify the cut-off frequency of the filter applied during demixing.
[0193] i. Modify the weighting applied to each microphone in the microphone array during demixing.
[0194] j. Modify the number of microphones in the microphone array.
[0195] k. Use the golden section search to determine the modified value of at least one parameter.
[0196] l. Use the Nelder-Mead algorithm to determine the modified value of at least one parameter.
[0197] m. Use the grid search to determine the modified value of at least one parameter.
[0198] n. Determine the modified value of one parameter of at least one parameter based on a predefined mapping between the error and at least one parameter.
[0199] o. Determine the mapping between the real-time determination error and at least one parameter.
[0200] In the foregoing specification, the present invention has been described with reference to specific examples of embodiments of the invention. However, it is apparent that various modifications and changes can be made herein without departing from the broader spirit and scope of the invention as set forth in the appended claims.
[0201] Furthermore, the terms "front", "rear", "top", "bottom", "above", "below", etc. in the specification and claims, if any, are used for descriptive purposes and do not necessarily describe a permanent relative position. It should be understood that such terms are interchangeable under appropriate circumstances, such that the embodiments of the present invention described herein can, for example, operate in other orientations different from those illustrated or otherwise described herein.
[0202] Any arrangement of components that achieves the same function is effectively "associated" such that the desired function is achieved. Thus, any two components that are combined herein to achieve a particular function can be regarded as being "associated" with each other such that the desired function is achieved, regardless of the architecture or intermediate components. Similarly, any two such associated components can also be regarded as being "operably connected" or "operably coupled" to each other to achieve the desired function.
[0203] Furthermore, those skilled in the art will recognize that the boundaries between the operations described above are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed among additional operations, and operations can be performed at least partially overlapping in time. Additionally, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in other embodiments.
[0204] However, other modifications, variations, and alternatives are also possible. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0205] The phrase "may be X" indicates that the condition X may be satisfied. This statement also indicates that the condition X may not be satisfied. For example - any reference to a system that includes a certain component should also cover the scenario where the system does not include that certain component. For example - any reference to a method that includes a certain step should also cover the scenario where the method does not include that certain component. For another example - any reference to a system configured to perform a certain operation should also cover the scenario where the system is not configured to perform that certain operation.
[0206] The terms "comprising", "including", "having", "consisting of", and "consisting essentially of" are used interchangeably. For example - any method can at least include the steps included in the drawings and / or the specification, and only include the steps included in the drawings and / or the specification. The same applies to systems.
[0207] The system can include a microphone array, a memory unit, and one or more hardware processors, such as a digital signal processor, FPGA, ASIC, general-purpose processor, etc. programmed to execute any of the methods mentioned above. The system can exclude the microphone array and instead can be fed from a sound signal generated by a microphone array.
[0208] It should be understood that, for simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, for clarity, the dimensions of some elements may be exaggerated relative to other elements. Additionally, where considered appropriate, reference numerals may be repeated within the drawings to indicate corresponding or similar elements.
[0209] In the foregoing specification, the invention has been described with reference to specific examples of embodiments of the invention. However, it is apparent that various modifications and changes can be made herein without departing from the broader spirit and scope of the invention as set forth in the appended claims.
[0210] Furthermore, the terms "front", "rear", "top", "bottom", "above", "below", etc. in the specification and claims, if any, are used for descriptive purposes and are not necessarily used to describe permanent relative positions. It should be understood that such terms are interchangeable where appropriate, such that the embodiments of the invention described herein can, for example, operate in other orientations different from those illustrated or otherwise described herein.
[0211] Those skilled in the art will recognize that the boundaries between logic blocks are merely illustrative, and alternative embodiments may combine logic blocks or circuit elements or impose alternative decompositions of the functionality on various logic blocks or circuit elements. Thus, it should be understood that the architecture described herein is merely exemplary, and in fact, many other architectures can be implemented that achieve the same functionality.
[0212] Any arrangement of components that achieves the same functionality is effectively "associated" such that the desired functionality is achieved. Thus, any two components combined herein to achieve a particular functionality can be seen as being "associated" with each other such that the desired functionality is achieved, regardless of the architecture or intermediate components. Similarly, any two such associated components can also be regarded as being "operably connected" or "operably coupled" to each other to achieve the desired functionality.
[0213] In addition, those skilled in the art will recognize that the boundaries between the operations described above are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed over additional operations, and operations can be performed at least partially overlapping in time. In addition, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in various other embodiments.
[0214] For example, in one embodiment, the illustrated examples can be implemented as circuitry located on a single integrated circuit or within the same device. Alternatively, the examples can be implemented as any number of separate integrated circuits or separate devices interconnected in a suitable manner.
[0215] For example, an example or portions thereof can be implemented as a physical circuit system or a soft or code representation that can be converted into a logical representation of a physical circuit system, such as in any suitable type of hardware description language.
[0216] In addition, the present invention is not limited to physical devices or units implemented in non-programmable hardware, but can be applied to programmable devices or units that perform the desired device functions by operating according to suitable program code, such as mainframes, minicomputers, servers, workstations, personal computers, laptop computers, personal digital assistants, electronic gaming machines, automobiles and other embedded systems, cellular telephones, and various other wireless devices, which are generally represented as "computer systems" in this application.
[0217] However, other modifications, variations, and alternatives are also possible. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0218] In a claim, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of other elements or steps than those listed in a claim. Further, as used herein, the term "a" or "an" is defined as one or more than one. Also, the introductory phrases such as "at least one" and "one or more" used in the claims shall not be construed to imply that the introduction of another claim element by the indefinite article "a" or "an" limits any particular claim containing such introduced claim element to inventions containing only one such element, even if the same claim includes the introductory phrases "one or more" or "at least one" as well as the indefinite article such as "a" or "an". The same holds for the use of definite articles. Terms such as "first" and "second" are used to arbitrarily distinguish between the elements so described. Thus, these terms are not necessarily intended to denote a temporal or other precedence of such elements. The fact that certain measures are recited in mutually different claims does not indicate that a combination of these measures cannot be used to advantage.
[0219] The invention may also be implemented in a computer program for running on a computer system, the computer program comprising at least code portions for performing the steps of a method according to the invention or for causing a programmable device to be able to perform the functions of a device or a system according to the invention when run on a programmable device such as a computer system. The computer program may cause a storage system to allocate disk drives to a disk drive group.
[0220] A computer program is a list of instructions such as a particular application program and / or an operating system. A computer program may for example comprise one or more of the following: subroutines, functions, programs, object methods, object implementations, executable applications, applets, servlets, source code, object code, shared libraries / dynamic loading libraries and / or other sequences of instructions designed to be executed on a computer system.
[0221] A computer program can be stored internally on a non-transitory computer-readable medium. All or some of the computer programs can be provided on a computer-readable medium that is permanently, removably, or remotely coupled to the information processing system. The computer-readable medium can include, for example, but is not limited to any number of the following: magnetic storage media, which includes disk and tape storage media; optical storage media, such as optical disc media (e.g., CD ROM, CD-R, etc.) and digital video disk storage media; non-volatile storage media, including semiconductor-based memory cells, such as flash memory, EEPROM, EPROM, ROM; ferromagnetic digital memories; MRAM; volatile storage media, including registers, buffers, or caches, main memory, RAM, etc. A computer process typically includes executing (running) a program or a part of a program, the current program values and state information, and the operating system to manage the resources used for the execution of the process. The operating system (OS) is software that manages the sharing of the resources of the computer and provides an interface for programmers to access these resources. The operating system processes system data and user input and responds by allocating and managing tasks and internal system resources as a service to the users and programs of the system. A computer system can include, for example, at least one processing unit, associated memory, and a plurality of input / output (I / O) devices. When a computer program is executed, the computer system processes information according to the computer program and generates the resulting output information through the I / O devices.
[0222] Any system mentioned in this patent application includes at least one hardware component.
[0223] Although certain features of the present invention have been illustrated and described herein, many modifications, substitutions, variations, and equivalents will now occur to those of ordinary skill in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and variations that fall within the true spirit of the present invention.
Claims
1. A source separation method, characterized in that, the method comprises: receiving or calculating an error related to a speech recognition process, the speech recognition process being applied to a previous speech recognition input based on a previous output of a source separation process; receiving a signal representing an audio signal originating from multiple sources and detected by multiple microphones; applying the source separation process to the signal based on the error to provide a source separation signal corresponding to audio signals of different sources originating from the multiple sources; and providing an output based on the source separation signal.
2. The method according to claim 1, comprising updating at least one parameter of the source separation process based on the error.
3. The method according to claim 2, comprising determining at least one updated parameter of the source separation process by updating the at least one parameter based on the error, and applying the source separation process to the signal according to the at least one updated parameter.
4. The method according to claim 2, comprising updating at least one parameter of the source separation process based on a change in the error.
5. The method according to claim 2, comprising updating at least one parameter of the source separation process based on a real-time change in the error.
6. The method according to claim 2, wherein applying the source separation process comprises applying a frequency conversion to convert the signal into multiple segments in the frequency domain, demixing the multiple segments to provide multiple demixed segments, and applying an inverse frequency conversion to convert the multiple demixed segments into the time domain, wherein, updating at least one parameter of the source separation process comprises updating at least one parameter of the frequency conversion and / or at least one parameter of the demixing.
7. The method according to claim 6, wherein updating at least one parameter of the source separation process comprises updating at least one of a segment length of the frequency conversion, a segment overlap of the frequency conversion, a sampling rate of the frequency conversion, and a window parameter of the frequency conversion.
8. The method according to claim 6, wherein updating at least one parameter of the source separation process comprises updating a filter cut-off frequency for the demixing, a microphone weighting for the demixing, and / or a number of microphones among the multiple microphones for the demixing.
9. The method according to claim 2, wherein updating at least one parameter of the source separation process comprises using at least one of a golden section search, a Nedler-Mead algorithm, and a grid search to adjust the at least one parameter.
10. The method according to claim 2, wherein updating at least one parameter of the source separation process comprises adjusting the at least one parameter based on a predefined mapping between the error and the at least one parameter.
11. The method according to claim 1, comprising providing the multiple source separation signals to a source selection process, wherein, the error is related to the speech recognition process, the speech recognition process being applied to a previous output of the source selection process.
12. The method according to claim 1, comprising determining the error based on an output from the speech recognition process.
13. The method according to claim 12, wherein the output from the speech recognition process includes the output text of the speech recognition process.
14. The method according to claim 13, comprising determining the error based on a comparison between the output text of the speech recognition process and a reference text.
15. The method according to claim 12, wherein the output of the speech recognition process includes a confidence score, the confidence score representing a reliability level of the output text of the speech recognition process.
16. The method according to claim 1, comprising receiving the error.
17. A source separator, comprising one or more processors, wherein, the one or more processors are configured to perform the method according to any one of claims 1 to 16.
18. A non - transitory computer - readable medium storing instructions, wherein, the instructions, when executed by a computerized system, cause the computerized system to perform the method according to any one of claims 1 to 16.
19. A speech recognition system, wherein, the system comprises: a plurality of microphones; a speech recognition engine that performs a speech recognition process; and a source separator configured to perform the method according to any one of claims 1 to 16.