Modulation-Domain Attention-Based Robust Voice Activity Detection in Reverberation and Noise
The method uses modulation frequency domain analysis to enhance speech detection by filtering reverberation and noise, addressing the challenges of noise and reverberation in audio systems, thereby improving speech quality and intelligibility.
Patent Information
- Application Number
- JP2024508558
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-28
- Filing Date
- 2022-08-11
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2042-08-11
AI Technical Summary
Existing speech enhancement systems struggle to adequately address noise and reverberation, leading to perceptual disruptions in audio/video content recording and playback systems.
A computer-implemented method for detecting speech from a reverberant signal using modulation frequency domain analysis, involving spectrotemporal amplitude enhancement, diffuseness index calculation, and feature extraction to filter out reverberation and noise.
Enhances speech detection accuracy by distinguishing between speech and noise, improving speech quality and intelligibility in noisy and reverberant environments.
Smart Images

Figure 0007791984000022 
Figure 0007791984000023 
Figure 0007791984000024
Abstract
Description
[Technical Field]
[0001] (Reference to Related Application) This application claims priority to the following priority applications: International Application No. PCT / CN2021 / 112265 (Reference No.: D20109WO), filed August 12, 2021; U.S. Provisional Application No. 63 / 239,976 (Reference No.: D20109USP1), filed September 2, 2021; and European Application No. EP21205203.9, filed October 28, 2021.
[0002] This application relates to voice activity detection. More particularly, the exemplary embodiments described below relate to solving noise and reverberation robustness problems based on modulation domain attention. [Background technology]
[0003] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued, and thus, unless otherwise indicated, it should not be assumed that any approach described in this section qualifies as prior art merely by virtue of its description in this section.
[0004] Traditionally, speech enhancement systems have been integrated into hands-free phones, video conferencing, or hearing aids, but have struggled to adequately address noise and reverberation (which may be considered noise but are referred to separately below). A robust voice activity detection (VAD) that estimates information about noise and reverberation and reduces artifacts and perceptual disruptions caused by noise and reverberation during speech would be useful. Such a VAD could be particularly useful for audio / video content recording and playback systems, such as the voice messaging component of any social networking software, video blog (vlog) platforms, or podcast setups, to improve speech quality and intelligibility. Summary of the Invention
[0005] A computer-implemented method for detecting speech from a reverberant signal based on data in the modulation frequency domain is disclosed, the method comprising the steps of: receiving new audio data in the time domain by a processor; converting, by the processor, one piece of the new audio data corresponding to a time point into a specific Spectrotemporal Amplitude (STA) as a time-frequency representation; applying a detection model to the specific STA to obtain an estimate of the degree of speech in the new audio data, wherein modulation spectral measurements (MSMs) having acoustic band dimensions and modulation band dimensions for the time point are obtained from one or more STAs obtained from the new audio data; calculating a diffuseness index (DI) indicating a degree of diffuseness in the modulation frequency domain for the piece of new audio data based on the MSMs; generating enhanced STAs by filtering reverberation and other noise from the specific STAs; creating one or more feature vectors using the DIs and the one or more features; determining an estimate of the degree of speech in the piece of new speech data from the one or more feature vectors; and transmitting the estimate of the degree of speech in the piece of new speech data. [Brief explanation of the drawings]
[0006] In the figures of the accompanying drawings, exemplary embodiments of the present invention are illustrated, without limitation, in which like reference numerals refer to similar elements and in which:
[0007] [Figure 1] FIG. 1 illustrates an exemplary networked computer system in which various embodiments may be implemented.
[0008] [Figure 2] FIG. 2 illustrates example components of an audio management server computer according to a disclosed embodiment.
[0009] [Figure 3A] FIG. 3A illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant audio signal with a reverberation time of 0 ms.
[0010] [Figure 3B] FIG. 3B illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant speech signal with a reverberation time of 500 ms.
[0011] [Figure 3C] FIG. 3C illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant speech signal with a reverberation time of 1 second (s).
[0012] [Figure 4A] FIG. 4A illustrates an energy plot in the acoustic / modulation frequency representation for noise recorded in a room, where the modulation frequency ranges up to 24 Hz.
[0013] [Figure 4B] Figure 4B illustrates an energy plot in the acoustic / modulation frequency representation for noise recorded in a room, where the modulation frequencies range from 4 to 24 Hz.
[0014] [Figure 5A] FIG. 5A illustrates an energy plot in the acoustic / modulation frequency representation with a signal-to-noise (SNR) ratio of 20 dB.
[0015] [Figure 5B] FIG. 5B illustrates an energy plot in the acoustic / modulation frequency representation with a signal-to-noise (SNR) ratio of 10 dB.
[0016] [Figure 5C] FIG. 5C illustrates an energy plot in the acoustic / modulation frequency representation where the signal-to-noise (SNR) ratio is 0 dB.
[0017] [Figure 6] FIG. 6 illustrates the process of enhancing the time-spectro-amplitude data in a spectro-temporal amplitude enhancer, with noise reduction performed by the audio management server computer.
[0018] [Figure 7] FIG. 7 illustrates an example process performed using an audio management server computer according to some embodiments described herein.
[0019] [Figure 8] FIG. 8 is a block diagram illustrating a computer system upon which one embodiment of the present invention may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0020] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of example embodiments of the present invention. However, it will be apparent that example embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the example embodiments.
[0021] In the following sections, embodiments are described according to the following outline: 1. Overview 2. Example of a Computing Environment 3. Examples of Computer Components 4. Functionality Description 4.1. Diffusion Indicator Module 4.2. Spectrotemporal Amplitude Enhancer 4.3. Enhancement Feature Extractor 4.4. Feature Fusion and Classification 5. Process Example 6. Hardware Implementation
[0022] 1. Overview A system and related methods are disclosed for detecting speech from a reverberant signal based on data in the modulation frequency domain. In some embodiments, the system is programmed to receive spectrotemporal amplitude data. The system is then programmed to enhance the spectrotemporal amplitude data by reducing reverberation and other noise and smoothing based on specific properties in the modulation frequency domain of a spectrotemporal spectrogram associated with the spectrotemporal amplitude data. The system is then programmed to calculate various features related to the presence of speech based on the enhanced spectrotemporal amplitude data and other data in the modulation frequency domain or the (acoustic) frequency domain. The system is then programmed to determine the extent of speech present in audio data corresponding to the received spectrotemporal amplitude data based on the various features. The system is programmable to transmit the extent of speech present to an output device.
[0023] In some embodiments, reducing reverberation to produce enhanced spectrotemporal amplitude data is primarily based on filtering out information within a specific modulation frequency range. Computing features that characterize the reduced presence of reverberation can involve applying existing metrics, typically applied to the frequency domain, to the modulation frequency domain, or directly extracting features from the modulation spectrogram related to the spectrotemporal amplitude.
[0024] The present system provides a technical advantage. It enables effective VAD by intelligently selecting features from speech data that distinguish between speech and noise (including reverberation). These features can be present at different levels, some related to environmental noise and some related to clean speech, and can be used to improve classification accuracy. Such VAD also enables detection and extraction of clean speech from given speech data, which has many applications, especially in environments where reverberation is common.
[0025] 2. Example of a Computing Environment Figure 1 illustrates an example of a networked computer system in which various embodiments may be implemented. Figure 1 is shown in simplified schematic form to provide a clear example; other embodiments may include more, fewer, or different elements.
[0026] In some embodiments, the networked computer system includes an audio management server computer 102 (“server”), one or more sensors 104 or input devices, and one or more output devices 110, which are communicatively connected via direct physical connections or via one or more networks 118.
[0027] In some embodiments, server 102 broadly represents one or more computers, virtual computing instances, and / or instances of applications programmed or configured with data structures and / or database records arranged to host or perform functions related to low-latency speech enhancement with noise reduction. Server 102 may comprise a server farm, a cloud computing platform, a parallel computer, or any other computing facility having sufficient computing power for data processing, data storage, and network communication for the functions described above.
[0028] In some embodiments, each of the one or more sensors 104 may include a microphone or other digital recording device that converts sound into an electrical signal. Each sensor is configured to transmit detected audio data to the server 102. Each sensor may include a processor or may be integrated into a typical client device, such as a desktop computer, laptop computer, tablet computer, smartphone, or wearable device.
[0029] In some embodiments, each of the one or more output devices 110 may include a speaker or other digital playback device that converts electrical signals back into sound. Each output device is programmed to play audio data received from the server 102. Like sensors, the output devices may include a processor or may be integrated into a typical client device, such as a desktop computer, laptop computer, tablet computer, smartphone, or wearable device.
[0030] The one or more networks 118 may be implemented by any medium or mechanism that provides for the exchange of data between the various elements of Figure 1. Examples of networks 118 include, but are not limited to, one or more of a cellular network communicatively connected to a data connection to a computing device via a cellular antenna, a near field communication (NFC) network, a local area network (LAN), a wide area network (WAN), the Internet, a terrestrial or satellite link, etc.
[0031] In some embodiments, the server 102 is programmed to receive input audio data corresponding to sounds in a given environment from one or more sensors 104. The server 102 is then programmed to process the input audio data (which typically represents a mixture of speech and noise) to estimate how much speech is present in each frame of the input data. The server 102 is also programmed to update the input audio data based on the estimate to generate cleaned-up output audio data that is expected to contain less noise than the input audio data. The server 102 is further programmed to send the output audio data to one or more output devices.
[0032] 3. Examples of Computer Components FIG. 2 illustrates example components of an audio management server computer according to disclosed embodiments. This diagram is for illustrative purposes only; the server 102 may include fewer or more functional or storage components. Each of the functional components may be implemented as a software component, a general-purpose or specialized hardware component, a firmware component, or any combination thereof. Each of the functional components may also be connected to one or more storage components (not shown). The storage components may be implemented using any of a relational database, an object database, a flat file system, or a JSON store. The storage components may connect to the functional components locally or over a network using program calls, a remote procedure call (RPC) facility, or a messaging bus. Components may or may not be self-contained. Depending on implementation-specific or other considerations, components may be functionally or physically centralized or distributed.
[0033] In some embodiments, the server 102 includes a modulation domain attention module 220. The modulation domain attention module 220 includes a diffuseness index module 202, a spectro-temporal amplitude enhancer 204, and an enhancement feature extractor 206. The server 102 also includes a feature fusion calculator 208 and a classification calculator 210.
[0034] In some embodiments, the diffuseness index module 202 includes computer-executable instructions that enable generation of distinguishable features that distinguish between speech and non-speech (e.g., reverberation or other noise) based on different clustering characteristics in the modulation frequency domain.
[0035] In some embodiments, the spectro-temporal amplitude enhancer 204 includes computer-executable instructions that enable enhancement of spectro-temporal amplitude in the modulation frequency domain for enhanced feature extraction.
[0036] In some embodiments, the enhancement feature extractor 206 includes computer-executable instructions that enable extraction of temporal and spectral features from the enhanced spectro-temporal amplitude data.
[0037] In some embodiments, the feature fusion operator 208 includes computer-executable instructions that enable the combination of features generated by the diffuseness index module 202, the enhancement feature extractor 206, and, optionally, other features as further described below.
[0038] In some embodiments, the classification calculator 210 includes computer-executable instructions that enable it to determine the presence of clean speech, free from reverberation or other noise, in given audio data based on a combination of features generated by the feature fusion calculator 208.
[0039] 4. Functionality Although mixed audio signals may have many overlaps in the time domain, modulation frequency analysis provides an additional dimension that may exhibit a higher degree of separation between audio sources. In other words, audio signals first captured in the time domain can be converted into a time-frequency representation (TFR) (a view of the signal as a function of time represented in both time and frequency) by a transform such as the discrete short-time Fourier transform (STFT). The TFR can then be expanded to a third dimension representing modulation frequency under given assumptions.
[0040] The modulation frequency domain is typically depicted by a modulation spectrogram, which shows intensity values, where the vertical axis represents the normal acoustic frequency index k and the horizontal axis represents the modulation frequency index i, as illustrated in Figures 3A, 3B, 4A, 4B, 5A, 5B, and 5C. Modulation spectrograms can show a greater degree of separation between audio sources.
[0041] In the modulation frequency domain, the temporal envelope of clean (anechoic) speech contains frequencies ranging from 2 to 16 Hz, with a spectral peak at approximately 4 Hz, corresponding to the syllabic rate of speech. However, noise and reverberation exhibit different modulation characteristics. In reverberant speech, the diffuse reverberation tail is often modeled as an exponentially decaying Gaussian white noise process. As the reverberation level increases, the signal acquires more Gaussian white noise-like properties. Reverberant signals exhibit a higher frequency temporal envelope due to the "whitening" effect of the reverberation tail.
[0042] FIG. 3A illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant speech signal with a reverberation time of 0 ms. FIG. 3B illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant speech signal with a reverberation time of 500 ms. FIG. 3C illustrates an energy plot in the acoustic / modulation frequency representation for a clean reverberant speech signal with a reverberation time of 1 second (s). As illustrated in FIG. 3B, for the clean speech, most of the modulation energy 302 is located primarily below 10 Hz in the modulation frequency domain, peaking at approximately 4 Hz. As illustrated in FIG. 3C, reverberation causes energy smearing to higher modulation frequencies. The stronger the reverberation, the greater the shift to higher modulation frequencies. These figures show that clean speech generally has higher energy, but concentrated in the lower modulation frequency domain, while the more reverberant the speech, the more energy shifts to the higher modulation frequency domain.
[0043] Because room noise diffusion typically occurs slowly as a function of time, the modulation spectrum of room noise is dominated by modulation frequencies below 1 Hz. Therefore, the envelope of room noise can be modeled as a constant plus random values. The constant envelope covers the main energy, concentrated below 1 Hz at the modulation frequency, and the random envelope covers the remaining energy, which is uniformly distributed across the entire modulation frequency range.
[0044] Figure 4A illustrates an energy plot in the acoustic / modulation frequency representation for noise recorded in a room (room noise without speech), where the modulation frequency ranges up to 24 Hz and is normalized from 0 to 24 Hz. As shown in Figure 4A, the main energy is concentrated below 1 Hz at the modulation frequency, illustrating the constant envelope of the room noise. Figure 4B illustrates a normalized energy plot in the acoustic / modulation frequency representation for noise recorded in a room, where the modulation frequency ranges from 4 to 24 Hz and is normalized from 4 to 24 Hz. As shown in Figure 4B, the residual random envelope of the room noise exhibits a uniform distribution along the modulation frequency dimension. Moreover, the energy is primarily concentrated at low acoustic frequencies and gradually decreases as the acoustic frequency increases.
[0045] FIG. 5A illustrates an energy plot in the acoustic / modulation frequency representation with a signal-to-noise (SNR) ratio of 20 dB. FIG. 5B illustrates an energy plot in the acoustic / modulation frequency representation with a signal-to-noise (SNR) ratio of 10 dB. FIG. 5C illustrates an energy plot in the acoustic / modulation frequency representation with a signal-to-noise (SNR) ratio of 0 dB. Here, the noise is the actual recorded room noise. As can be seen from these figures, when the modulation frequency is above 4 Hz, the constant time envelope of the room noise shown in FIG. 4A (below 4 Hz) is filtered or masked, while the remaining random time envelope of the noise shown in FIG. 4B masks the speech region relatively uniformly, especially in the low acoustic frequency band. The stronger the noise, the greater the masking in the modulation frequency region. This can be seen from the fact that the proportion of high-energy portions, such as 502, within most portions of the same acoustic frequency band, such as 504 in FIG. 5C, is smaller than the proportion of high-energy portions, such as 506, within most portions of the same acoustic frequency band. Therefore, most of the energy is present in the low acoustic frequency range. In addition, clean speech generally produces high energy, but concentrated in the lower modulation frequency range, and the more "noisy" the sound, the greater the blending or masking of random time envelopes that introduce energy into the higher modulation frequency range.
[0046] In some embodiments, the server 102 receives a time-domain signal x(n), where n represents a discrete time-dependent variable. The time-frequency (TF) transform X(l,k) of x(n) can be obtained using the STFT.
number
[0047] In some embodiments, the server 102 then transforms the TF transformed narrowband signal X(l,k) into spectrotemporal amplitudes Y(l,m) of the perceptual acoustic band based on the human auditory system using the following transformation matrix:
number
[0048] In some embodiments, the modulation spectral measurement (spectrogram) Z(l,m,c) at any frame l, perceptual acoustic band m, and modulation band c is calculated using the last L frames of the spectral amplitude based FFT.
number
[0049] 4.1. Diffusion Indicator Module In some embodiments, the server 102 in the diffuseness module 202 calculates a diffuseness indicator (DI) based on the last L frames, characterizing the relationship between the energy in the lower range of the modulation frequency domain and the energy in the higher range of the modulation frequency domain for a particular time. As mentioned above, energy data corresponding to clean speech tends to be in the lower range of the modulation frequency domain, but the more reverberation and other noise that is mixed with clean speech, the more the energy data corresponding to that mixture tends to spread out into the higher range of the modulation frequency domain, resulting in a greater "spread" of energy values in the modulation frequency domain. Thus, a higher DI indicates a more reverberant or otherwise noisy audio signal.
[0050] In some embodiments, the DI can be calculated as the centroid of the modulation spectrum.
number
[0051] In some embodiments, the DI can be calculated as the energy ratio between the low modulation portion and the high modulation portion.
number
[0052] In some embodiments, the diffuseness index can be calculated as the energy ratio of the low modulation portion to the full modulation portion.
number
[0053] 4.2. Spectrotemporal Amplitude Enhancer 6 illustrates a process of enhancing spectro-temporal amplitude data in a spectro-temporal amplitude enhancer, with noise reduction performed by the server. In some embodiments, the server 102 in the spectro-temporal amplitude enhancer 204 performs a series of steps, including reverberation and noise filtering in the modulation frequency domain, residual noise estimation, and residual noise suppression, to convert the initial spectro-temporal amplitude data into enhanced spectro-temporal amplitude data.
[0054] In some embodiments, given the modulation spectrum measurements calculated from equation (1), the server 102 in box 604 filters noise and reverberation to generate filtered modulation spectrum measurements.
number
number
[0055] In some embodiments, the server 102 in box 606 smooths the filtered modulation spectrum measurements as follows: Parseval's theorem, which roughly states that the sum of the squares (or integral) of a function is equal to the sum of the squares (or integral) of its Fourier transform:
number
[0056] The server 102 calculates the smoothed spectrotemporal energy in the modulation frequency domain by summing
number
number
[0057] Here, the server 102 calculates the enhanced spectrotemporal energy (EME) with reverberation and noise filtering in the modulation frequency domain based on equations (5) and (6) above.
number
number
[0058] Next, the server 102 calculates the smoothed emphasized spectral temporal amplitude based on the above equation (7).
number
number
[0059] In some embodiments, the server 102 in box 608 measures the spectrotemporal amplitude of the residual (ambient) noise.
number
[0060] In some embodiments, the server 102 in box 610 performs residual noise estimation and suppression to generate enhanced spectrotemporal amplitudes as output data in box 620 as follows:
number
number
[0061] In some embodiments, data in the modulation frequency domain can be used to calculate enhanced spectrotemporal amplitudes via a machine learning model. To build such a model, a training dataset can include, as input data, an "original speech" class containing several spectrotemporal amplitude data in the modulation frequency domain corresponding to combinations of clean speech, noise, and reverberation over a range of lengths (e.g., 5 minutes). The training dataset can include, as output data, an "enhanced speech" class containing several spectrotemporal amplitude data in the modulation frequency domain corresponding to clean speech that has been smoothed and noise-reduced. As described above, the noise reduction includes removal of reverberation, ambient sounds, and other noises. Machine learning methods known to those skilled in the art, such as those described in arXiv:1709.08243 or arXiv:1704.07804 [cs.CV], can then be applied to the training dataset to build a model configured to generate enhanced spectrotemporal amplitude data. A feature extractor can then extract features based on the enhanced spectrotemporal amplitudes instead of the original amplitudes to derive the enhanced features, as described below.
[0062] 4.3. Enhancement Feature Extractor In some embodiments, the server 102 in the enhancement feature extractor 206 calculates specific features of the enhanced temporal spectral amplitude, such as enhanced Mel-Frequency Cepstral Coefficients (MFCCs) or enhanced Spectral Flatness (SFT), which are typically applied to frequency spectra.
[0063] In some embodiments, the server 102 calculates emphasized MFCCs (EMFCCs) by using the emphasized time-spectral amplitudes calculated in the spectro-temporal amplitude emphasizer 204 instead of the original spectro-temporal amplitudes in calculating the MFCCs. The Mel-frequency filter can be treated as a specific banding matrix before calculating the MFCCs.
[0064] In some embodiments, the server 102 calculates an enhanced SFT (ESFT) by using the enhanced spectro-temporal amplitude calculated in the spectro-temporal amplitude enhancer 204 instead of the original spectro-temporal amplitude in calculating the SFT. Specifically, the original SFT can be calculated as follows, using Y(l,m) to account for the time dimension:
number
number
number
[0065] In some embodiments, other spectrally related measurements can also be used to characterize the flatness or peakiness of the signal spectrum, or to generate further features of the enhanced spectrotemporal amplitude, such as: ●Spectral peak based on the sum of the power ratios of the peak band and other bands Spectral peak based on the power ratio between the peak and the average (no peak band) Variance or standard deviation of adjacent spectral band power The sum or maximum value of the spectral band power difference between adjacent frequency bands Spectral spread or dispersion around the spectral center Spectral entropy
[0066] 4.4. Feature Fusion and Classification In some embodiments, the server 102 in the feature fusion calculator 208 combines the diffuseness index, the enhancement features, and other commonly used features without enhancement (such as zero-crossing rate in the frequency domain, spectral flux, or pitch). The server 102 then calculates one or more feature vectors from the combination. The outputs of all the features may be simply concatenated into one feature to form a one-feature vector. Alternatively, different features may form one multi-feature vector. Alternatively, different features may form respective feature vectors, each vector having one feature.
[0067] In some embodiments, the server 102 in the classification calculator 210 classifies one or more feature vectors generated by the feature fusion calculator 208 through a machine learning model. To build the model, the server 102 may prepare a training set of feature vectors generated by applying a set of audio signals (transformed into the frequency domain and modulation frequency domain) containing varying degrees of speech (without reverberation or other noise) and varying degrees of reverberation to the modules 202, 204, 206, and 208. "Degree" may be defined as a proportion of volume or loudness, i.e., the amplitude of a sound wave, or another sound characteristic. For each signal in the training set, the input data may be the extracted feature vector, and the output data may be an indication of the presence or absence of any speech in the signal (a binary value) or the degree of clean speech in the signal (a continuous value). The server 102 can then apply any machine learning model for classification known to those skilled in the art, such as statistical methods including logistic regression, adaptive boosting (AdaBoost), or Gaussian mixture models (GMM), or artificial neural networks including multi-layer perceptrons or support vector machines. For example, in the case of neural networks, a softmax function can be applied to calculate the probability that the input signal contains speech. This probability can be used as an estimate of the degree of speech in the input signal.
[0068] 5. Processing Examples FIG. 7 illustrates an example process performed using an audio management server computer according to some embodiments described herein. FIG. 7 is shown in a simplified schematic format for purposes of illustrating a clear example; other embodiments may include more, fewer, or different elements connected in various ways. FIG. 7 is intended to disclose an algorithm, plan, or outline that can be used to implement one or more computer programs or other software elements that, when executed, effect the functional improvements and technical advances described herein. Furthermore, the flow diagrams herein are described with the same level of detail that those skilled in the art would typically use to communicate with one another, using their accumulated skill and knowledge, about the algorithms, plans, or specifications that form the basis of the software programs they plan to code or implement.
[0069] In some embodiments, in step 702, the server 102 is programmed to receive new audio data in the time domain.
[0070] In some embodiments, in step 704, the server 102 is programmed to convert the new audio data corresponding to a point in time into a particular Spectrotemporal Amplitude (STA) time-frequency representation.
[0071] In some embodiments, in step 706, the server 102 is programmed to obtain modulation spectrum measurements (MSMs) having acoustic band dimensions and modulation band dimensions for the above-mentioned time point from one or more STAs obtained from the new audio data.
[0072] In some embodiments, in step 708, the server 102 is programmed to calculate a diffuseness index (DI) indicating the degree of diffuseness in the modulation frequency domain for the piece of new audio data based on the MSM.
[0073] In some embodiments, the DI is the center of gravity of the modulation spectrum based on the values of MSM in a range of the modulation frequency band and a range of the audio frequency band. In other embodiments, the DI is the energy ratio between a low modulation portion based on the values of MSM in a low range of the modulation frequency band and a range of the audio frequency band, and a high modulation portion based on the values of MSM in a high range of the modulation frequency band and a range of the audio frequency band. In other embodiments, the DI is the energy ratio between a low modulation portion based on the values of MSM in a low range of the modulation frequency band and a range of the audio frequency band, and a full modulation portion based on the values of MSM in the entire range of the modulation frequency band and that range of the audio frequency band.
[0074] In some embodiments, calculating the DI involves applying a machine learning model trained with measurements of MSM for audio data having only clean speech and audio data having different degrees of reverberation and other noise as input data, and corresponding DI values as output data.
[0075] In some embodiments, in step 710, the server 102 is programmed to generate an enhanced STA that has filtered reverberation and other noise from a particular STA.
[0076] In some embodiments, generating the enhanced STA includes filtering out values of the MSM outside a range of a modulation frequency band, which in other embodiments is between 3 and 30 Hz.
[0077] In some embodiments, generating the enhanced STAs includes calculating a smoothed spectrotemporal energy by aggregating over time, while in other embodiments, generating the enhanced STAs includes removing residual noise by tracking minimum spectrotemporal energy over time.
[0078] In some embodiments, generating the enhanced STA includes applying a machine learning model trained with spectro-temporal amplitude data corresponding to different degrees of reverberation and other noise as input data and corresponding spectro-temporal amplitude data corresponding to only clean speech as output data. In other embodiments, the server 102 is further programmed to extract features characterizing clean speech, including a low cutoff modulation frequency and a high cutoff modulation frequency, from the application of the machine learning model.
[0079] In some embodiments, in step 712, the server 102 is programmed to calculate one or more features from the enhanced STA and create one or more feature vectors using the DI and the one or more features.
[0080] In some embodiments, the calculating includes calculating enhanced Mel-frequency filtered cepstral coefficients (MFCCs), where the enhanced MFCCs are calculated by applying a Mel-frequency filter to the enhanced STAs for use in the final step of calculating the MFCCs. In other embodiments, the calculating includes calculating enhanced spectral flatness (SFT), where the enhanced SFT is calculated by using the enhanced STAs instead of the STAs and summing values over time in calculating the SFT.
[0081] In some embodiments, the one or more features include a spectral peak based on the sum of the power ratios of the peak band to other bands, a spectral peak based on the power ratio of the peak to the average (without the peak band), a variance or standard deviation of adjacent spectral band power, a sum or maximum of the spectral band power difference between adjacent frequency bands, a spectral spread or spectral dispersion around the spectral center, and a spectral entropy.
[0082] In some embodiments, in step 714, server 102 is programmed to determine an estimate of the degree of speech in the piece of new audio data from the one or more feature vectors and transmit the estimate of the degree of speech in the piece of new audio data.
[0083] In some embodiments, the determining includes applying a machine learning model trained with one or more features of spectro-temporal amplitude data corresponding to clean speech and spectro-temporal amplitude data corresponding to different degrees of reverberation and other noise as input data, and the corresponding speech degrees as output data.
[0084] 6. Hardware Implementation According to one embodiment, the techniques described herein are implemented by at least one computing device. The techniques may be implemented, in whole or in part, using a combination of at least one server computer and / or other computing devices connected using a network, such as a packet data network. The computing device may be hardwired to perform the techniques, or may include digital electronic devices such as at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA) persistently programmed to perform the techniques, or may include at least one general-purpose hardware processor programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination. Such computing devices may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to achieve the techniques. The computing device may be a server computer, a workstation, a personal computer, a portable computer system, a handheld device, a mobile computing device, a wearable device, a body-worn or implantable device, a smartphone, a smart appliance, an internetworking device, an autonomous or semi-autonomous device (such as a robot or an unmanned ground or air vehicle), any other electronic device embedded with hardwired and / or program logic to implement the above techniques, one or more virtual computing machines or instances in a data center, and / or a network of server computers and / or personal computers.
[0085] 8 is a block diagram illustrating an example of a computer system in which an embodiment may be implemented. In the example of FIG. 8, a computer system 800 and instructions for implementing the disclosed techniques in hardware, software, or a combination of hardware and software are represented diagrammatically, e.g., as boxes and circles, at a level of detail commonly used by those skilled in the art to which this disclosure pertains, to facilitate understanding of computer architecture and computer system implementation.
[0086] Computer system 800 includes an input / output (I / O) subsystem 802, which may include a bus and / or other communication mechanisms for communicating information and / or instructions between components of computer system 800 via electronic signal paths. I / O subsystem 802 may include an I / O controller, a memory controller, and at least one I / O port. The electronic signal paths are represented schematically in the figure, for example, as lines, single-headed arrows, or double-headed arrows.
[0087] At least one hardware processor 804 is connected to the I / O subsystem 802 for processing information and instructions. The hardware processor 804 may include, for example, a general-purpose microprocessor or microcontroller and / or a special-purpose microprocessor such as an embedded system, a graphics processing unit (GPU), a digital signal processor, or an ARM processor. The processor 804 may include an integrated arithmetic logic unit (ALU) or may be connected to a separate ALU.
[0088] Computer system 800 includes memory 806, consisting of one or more units (e.g., main memory), coupled to I / O subsystem 802 for electronically and digitally storing data and instructions to be executed by processor 804. Memory 806 may include volatile memory, such as various forms of random access memory (RAM) or other dynamic storage devices. Memory 806 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 804. Such instructions, when stored on a non-transitory computer-readable storage medium accessible to processor 804, may turn computer system 800 into a special-purpose machine customized to perform the operations specified in the instructions.
[0089] Computer system 800 further includes non-volatile memory (such as read-only memory (ROM) 808 or other static storage device) connected to I / O subsystem 802 for storing information and instructions for processor 804. ROM 808 may include various forms of programmable ROM (PROM), such as erasable programmable read-only memory (EPROM) or electrically erasable programmable read-only memory (EEPROM). A unit of persistent storage 810, which may include various forms of non-volatile random access memory (NVRAM), such as flash memory, or solid-state storage, a magnetic disk, or an optical disk, such as a CD-ROM or DVD-ROM, may be connected to I / O subsystem 802 for storing information and instructions. Storage 810 is one example of a non-transitory computer-readable medium that may be used to store instructions and data that, when executed by processor 804, cause computer-implemented methods to perform the techniques described herein.
[0090] The instructions in memory 806, ROM 808, or storage 810 may include one or more sets of instructions organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized as one or more computer programs, operating system services, or application programs, including mobile apps. The instructions may include operating systems and / or system software; one or more libraries supporting multimedia, programming, or other functionality; data protocol instructions or stacks implementing TCP / IP, HTTP, or other communications protocols; file processing instructions for interpreting and rendering files coded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command line interface, or text user interface; application software such as an office suite, Internet access application, design and manufacturing application, graphics application, audio application, software engineering application, educational application, game, or other application. The instructions may implement a web server, web application server, or web client. The instructions may be organized as a presentation layer, an application layer, and a data storage layer, such as a relational database system using Structured Query Language (SQL) or NoSQL, an object store, a graph database, a flat file system, or other data storage.
[0091] Computer system 800 may be connected to at least one output device 812 via I / O subsystem 802. In one embodiment, output device 812 is a digital computer display. Examples of displays that may be used in various embodiments include a touchscreen display or a light-emitting diode (LED) display or a liquid crystal display (LCD) or an electronic paper display. Computer system 800 may include other types of output device 812 instead of, or in addition to, a display device. Examples of other output device 812 include a printer, a ticket printer, a plotter, a projector, a sound or video card, a speaker, a buzzer or piezoelectric device or other audible device, a lamp or LED or LCD indicator, a tactile device, an actuator, or a servo.
[0092] At least one input device 814 is connected to the I / O subsystem 802 for communicating signals, data, command selections, or gestures to the processor 804. Examples of input device 814 include touch screens, microphones, still and video digital cameras, alphanumeric and other keys, keypads, keyboards, graphics tablets, image scanners, joysticks, clocks, switches, buttons, dials, slides, and / or various types of sensors such as force sensors, motion sensors, thermal sensors, accelerometers, gyroscopes, and inertial measurement unit (IMU) sensors, and / or various types of transceivers such as wireless, radio frequency (RF) or infrared (IR) transceivers such as cellular or Wi-Fi, and global positioning system (GPS) transceivers.
[0093] Another type of input device is the control device 816. The control device 816 may perform cursor control or other automated control functions, such as navigating a graphical interface on a display screen, instead of or in addition to input functions. The control device 816 may be a touchpad, mouse, trackball, or cursor direction keys for communicating directional information and command selections to the processor 804 and for controlling cursor movement on the display 812. The input device may have at least two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allow the device to locate a position in a plane. Another type of input device is a wired, wireless, or optical control device, such as a joystick, wand, console, steering wheel, pedals, gear shift mechanism, or other type of control device. The input device 814 may include a combination of multiple different input devices, such as a video camera and a depth sensor.
[0094] In another embodiment, computer system 800 may comprise an Internet of Things (IoT) device that omits one or more of output device(s) 812, input device(s) 814, and control device(s) 816. Alternatively, in such an embodiment, input device(s) 814 may comprise one or more cameras, motion detectors, thermometers, microphones, earthquake detectors, other sensors or detectors, measurement devices, or encoders, and output device(s) 812 may comprise a dedicated display, such as a single-line LED or LCD display, one or more indicators, a display panel, a meter, a valve, a solenoid, an actuator, or a servo.
[0095] If computer system 800 is a mobile computing device, input device 814 may include a Global Positioning System (GPS) receiver connected to a GPS module that can triangulate against multiple GPS satellites to determine and generate geolocation or position data, such as latitude-longitude values, for the geophysical location of computer system 800. Output device 812 may include hardware, software, firmware, and interfaces for generating position report packets, notifications, pulse or heartbeat signals, or other recurring data transmissions that identify the location of computer system 800, alone or in combination with other application-specific data, to host 824 or server 830.
[0096] Computer system 800 may implement the techniques described herein using customized hardwired logic, at least one ASIC or FPGA, firmware, and / or program instructions or logic that, when loaded and used or executed in conjunction with the computer system, cause or program the computer system to operate as a special-purpose machine. According to one embodiment, the techniques are performed by computer system 800 in response to processor 804 executing at least one sequence of at least one instruction contained in main memory 806. Such instructions may be read into main memory 806 from another storage medium, such as storage 810. Execution of the sequences of instructions contained in main memory 806 causes processor 804 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0097] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage 810. Volatile media include, for example, dynamic memory, such as memory 806. Common forms of storage media include, for example, hard disks, solid-state drives, flash drives, magnetic data storage media, any optical or physical data storage media, memory chips, etc.
[0098] Storage media are distinct from, but may be used in conjunction with, transmission media. Transmission media involves transferring information between storage media. For example, transmission media include coaxial cables, copper wire, and fiber optics, including wires such as a bus in I / O subsystem 802. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0099] Various forms of media may be involved in carrying at least one sequence of at least one instruction to processor 804 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into its own dynamic memory and transmit the instructions over a communications link, such as a modem, fiber optic or coaxial cable, or a telephone line. Data on the communications link may be received by a modem or router local to computer system 800 and converted so as to be readable by computer system 800. For example, data carried by a radio or optical signal may be received by a receiver, such as a radio frequency antenna or infrared detector, and provided by appropriate circuitry to I / O subsystem 802 (e.g., placing the data on a bus). I / O subsystem 802 carries the data to memory 806. The data is retrieved from memory 806 by processor 804, and the instructions are executed. The instructions received by memory 806 may optionally be stored on storage 810 either before or after execution by processor 804.
[0100] Computer system 800 also includes a communications interface 818 coupled to bus 802. Communications interface 818 provides a two-way data communication coupling to a network link 820 that is directly or indirectly connected to at least one communications network, such as a network 822 or a public or private cloud on the Internet. For example, communications interface 818 may be an Ethernet networking interface, an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of communications line, such as an Ethernet cable or any type of metallic or fiber optic line or a telephone line. Network 822 broadly represents a local area network (LAN), a wide area network (WAN), a campus network, an internetwork, or any combination thereof. Communications interface 818 may comprise a LAN card that provides a data communication connection to a compatible LAN, or a wired cellular radiotelephone interface for transmitting or receiving cellular data in accordance with a cellular radiotelephone wireless networking standard, or a wired satellite radio interface for transmitting or receiving digital data in accordance with a satellite wireless networking standard. In any such implementation, communication interface 818 sends and receives electrical, electromagnetic or optical signals over signal paths that carry digital data streams representing various types of information.
[0101] Network link 820 typically provides electrical, electromagnetic, or optical data communication to other data devices directly or through at least one network, for example using satellite, cellular, Wi-Fi, or Bluetooth technology. For example, network link 820 may provide a connection through network 822 to a host computer 824.
[0102] Further, network link 820 may provide connectivity through network 822 or to other computing devices through internetworking devices and / or computers operated by an Internet Service Provider (ISP) 826. ISP 826 provides data communication services through a worldwide packet data communication network represented as Internet 828. Connected to Internet 828 may be a server computer 830. Server 830 broadly represents any computer, data center, virtual machine or virtual computing instance with or without a hypervisor, or a computer running a containerized program system such as DOCKER or KUBERNETES. Server 830 may represent an electronic digital service implemented using two or more computers or instances and accessed and used by sending a web service request, a uniform resource locator (URL) string with parameters in an HTTP payload, an API call, an app service call, or other service call. Computer system 800 and server 830 may form elements of a distributed computing system that includes other computers, processing clusters, server farms, or other configurations of computers that cooperate to perform tasks or run applications or services. The server 830 may comprise one or more sets of instructions organized as modules, methods, objects, functions, routines, or calls, which may be organized as one or more computer programs, operating system services, or application programs, including mobile apps.The instructions may include operating system and / or system software; one or more libraries supporting multimedia, programming, or other functionality; data protocol instructions or stacks implementing TCP / IP, HTTP, or other communications protocols; file formatting instructions for interpreting or rendering files coded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command line interface, or text user interface; and application software such as an office suite, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games, or other applications. Server 830 may comprise a web application server hosting a presentation layer, an application layer, and a data storage layer such as a relational database system using Structured Query Language (SQL) or NoSQL, an object store, a graph database, a flat file system, or other data storage.
[0103] Computer system 800 can send messages and receive instructions, including data and program code, through the network(s), network link 820 and communication interface 818. In the Internet example, a server 830 might transmit a requested code for an application program through Internet 828, ISP 826, local network 822 and communication interface 818. The received code may be executed by processor 804 as it is received, and / or stored in storage 810, or other non-volatile storage for later execution.
[0104] Execution of the instructions described in this section may implement a process in the form of a running computer program instance, which consists of program code and its current operations. Depending on the operating system (OS), a process may consist of multiple threads of execution that execute instructions simultaneously. In this sense, a computer program is a passive collection of instructions, while a process may be the actual execution of those instructions. Multiple processes may relate to the same program. For example, opening multiple instances of the same program often means that two or more processes are running. Multitasking may be implemented to allow multiple processes to share the processor 804. Although each processor 804 or each core of that processor executes one task at a time, the computer system 800 may be programmed to implement multitasking to allow each processor to switch between multiple running tasks without having to wait for each task to finish. In one embodiment, switching may occur when a task performs an input / output operation, when the task indicates that it is available to switch, or upon a hardware interrupt. Time sharing may be implemented to enable fast response for interactive user applications by performing rapid context switching to allow multiple processes to appear to be running simultaneously. In one embodiment, for safety and reliability, the operating system may prevent direct communication between independent processes and provide a strictly mediated and controlled inter-process communication facility.
[0105] 7. Extensions and Substitutions In the above, the disclosed embodiments have been described with reference to many specific details that may vary from implementation to implementation. Thus, the specification and drawings should be considered in an illustrative, not a limiting sense. What is the sole and exclusive indication of the scope of the present disclosure, and what the applicants intend to be the scope of the present disclosure, is the literal and equivalent scope of the set of claims issuing from this application, including any subsequent amendments, in the specific form in which such claims arise.
[0106] Various aspects of the present invention can be understood from the following enumerated example descriptions (EEE). EEE1. 1. A computer-implemented method for detecting speech from a reverberant signal based on data in the modulation frequency domain, comprising: obtaining, by a processor, a specific Spectrotemporal Amplitude (STA) corresponding to a point in time covered by the new audio data in the time domain as a time-frequency representation; obtaining modulation spectral measurements (MSMs) for the time point from one or more STAs obtained from new audio data, the MSMs having acoustic band dimensions and modulation band dimensions; calculating a diffuseness index (DI) indicating a degree of diffuseness in a modulation frequency domain for the one new audio data based on the MSM; generating an enhanced STA from the specific STA by filtering out reverberation and other noise; computing one or more features from the enhanced STA; creating one or more feature vectors using the DI and the one or more features; determining an estimate of the degree of voice in the piece of new audio data from the one or more feature vectors; outputting the estimate of the degree of voice in the piece of new audio data; A computer-implemented method comprising: EEE2. The DI is a center of gravity of a modulation spectrum based on the values of the MSM in a range of modulation frequency bands and a range of acoustic frequency bands. A method performed by a computer in EEE1. EEE3. The DI is an energy ratio between a low modulation part based on the value of the MSM in a low range of the modulation frequency band and a range of the audio frequency band, and a high modulation part based on the value of the MSM in a high range of the modulation frequency band and the same range of the audio frequency band. A method performed by a computer in EEE1. EEE4. The DI is an energy ratio between a low modulation part based on the values of the MSM in a low range of a modulation frequency band and a range of an audio frequency band, and a full modulation part based on the values of the MSM in the entire range of the modulation frequency band and the said range of the audio frequency band. A method performed by a computer in EEE1. EEE5. the obtaining step includes calculating the MSM using a Fast Fourier Transform using a plurality of new audio data corresponding to a predetermined number of consecutive time points before the time point; A method performed by a computer in EEE1. EEE6. generating the enhanced STA includes filtering out values of the MSM outside an excluded range of a modulation frequency band; A method executed by any one of the computers EEE1 to EEE5. EEE7. The excluded range of the modulation frequency band is 3 to 30 Hz; A method performed by an EEE6 computer. EEE8. generating the enhanced STA includes calculating a smoothed spectro-temporal energy by aggregating over time; A method executed by any one of the computers EEE1 to EEE7. EEE9. generating the enhanced STA includes removing residual noise by tracking minimum spectro-temporal energy over time. A method executed by any one of the computers EEE1 to EEE8. EEE10. generating the enhanced STA includes applying a machine learning model trained using spectro-temporal amplitude data corresponding to different degrees of reverberation and other noise as input data and corresponding spectro-temporal amplitude data corresponding to only clean speech as output data; A method executed by any one of the computers EEE1 to EEE7. EEE11. extracting features characterizing the clean speech from the application of the machine learning model, including a low cutoff modulation frequency and a high cutoff modulation frequency; The computer-implemented method of EEE10 further comprising: EEE12. the calculating step includes calculating enhanced Mel-Frequency Cepstral Coefficients (MFCCs) using the enhanced STAs; A method executed by any one of the computers EEE1 to EEE11. EEE13. the calculating step includes calculating an enhanced spectral flatness (SFT), using the enhanced STA instead of the STA, and calculating the enhanced SFT by summing values over time in calculating the SFT. A method executed by any one of the computers EEE1 to EEE12. EEE14. The one or more features include a spectral peak based on the sum of the power ratios of the peak band to other bands, a spectral peak based on the power ratio of the peak to the average (without the peak band), a variance or standard deviation of adjacent spectral band power, a sum or maximum of the spectral band power difference between adjacent frequency bands, a spectral spread or spectral dispersion around a spectral center, and a spectral entropy. A method performed by any one of the computers EEE1 to EEE13. EEE15. the determining step includes applying a machine learning model trained using one or more features of spectro-temporal amplitude data corresponding to clean speech and spectro-temporal amplitude data corresponding to different degrees of reverberation and other noise as input data, and corresponding speech degrees as output data. A method performed by any one of the computers EEE1 to EEE14. EEE16. receiving new audio data in the time domain; converting the new audio data corresponding to a single time point into the specified Spectro-Temporal Amplitude (STA) as a time-frequency representation; The computer-implemented method of any one of EEE1 to 15, further comprising: EEE17. 1. A computer-implemented method for detecting speech from a reverberant signal based on data in the modulation frequency domain, comprising: obtaining, by a processor, new audio data in the time domain; converting one piece of said new audio data corresponding to a time point into a specific Spectro-Temporal Amplitude (STA) as a time-frequency representation; applying a detection model to the particular STA to obtain an estimate of the degree of voicedness in the new audio data; wherein the applying step comprises: obtaining, by the processor, modulation spectral measurements (MSMs) for the time point from one or more STAs obtained from new audio data, the MSMs having acoustic band dimensions and modulation band dimensions; Calculating a diffuseness index (DI) indicating a degree of diffuseness in a modulation frequency domain for one of the new audio data corresponding to the time point based on the MSM; generating an enhanced STA from the specific STA by filtering out reverberation and other noise; computing one or more features from the enhanced STA; creating one or more feature vectors using the DI and the one or more features; determining an estimate of the degree of voice in the piece of new audio data from the one or more feature vectors; outputting the estimate of the degree of voice in the piece of new audio data; Including, A computer-implemented method. EEE18. the obtaining step includes calculating the MSM using a Fast Fourier Transform using a plurality of new audio data corresponding to a predetermined number of consecutive time points before the time point; A method implemented by a computer in EEE17. EEE19. The generating step is based on Parseval's theorem. A method implemented by a computer in EEE17. EEE20. The calculating step includes using values of the MSM having a range of an acoustic frequency band of 125 to 8,000 Hz. A method implemented by a computer in EEE17.
Claims
1. 1. A computer-implemented method for detecting speech from a reverberant signal based on data in the modulation frequency domain, comprising: obtaining, by a processor, a specific Spectrotemporal Amplitude (STA) corresponding to a point in time covered by one new audio data in the time domain as a time-frequency representation; Obtaining modulation spectral measurements (MSMs) for the time point, having acoustic band dimensions and modulation band dimensions, from one or more STAs obtained from the piece of new audio data; calculating a diffuseness index (DI) indicating a degree of diffuseness in a modulation frequency domain for the new audio data based on the MSM; generating an enhanced STA from the particular STA by filtering out reverberation and other noise; Computing one or more features from the enhanced STA; creating one or more feature vectors using the DI and the one or more features; determining an estimate of the degree of voice in the piece of new audio data from the one or more feature vectors; outputting the estimate of the degree of voice in the piece of new audio data; A computer-implemented method comprising:
2. The DI is a center of gravity of a modulation spectrum based on the values of the MSM in a range of modulation frequency bands and a range of acoustic frequency bands. The computer-implemented method of claim 1.
3. The DI is an energy ratio between a low modulation part based on the value of the MSM in a low range of a modulation frequency band and a range of an audio frequency band, and a high modulation part based on the value of the MSM in a high range of a modulation frequency band and the same range of the audio frequency band. The computer-implemented method of claim 1.
4. The DI is an energy ratio between a low modulation part based on the values of the MSM in a low range of a modulation frequency band and a range of an audio frequency band, and a full modulation part based on the values of the MSM in the entire range of the modulation frequency band and the said range of the audio frequency band.
10. The computer-implemented method of claim 1.
5. the obtaining step includes calculating the MSM using a Fast Fourier Transform using a plurality of new audio data corresponding to a predetermined number of consecutive time points before the time point; The computer-implemented method of claim 1.
6. generating the enhanced STA includes filtering out values of the MSM outside an excluded range of a modulation frequency band; A computer-implemented method according to any one of claims 1 to 5.
7. the excluded range of the modulation frequency band is 3 to 30 Hz; 7. The computer-implemented method of claim 6.
8. generating the enhanced STA includes calculating a smoothed spectro-temporal energy by aggregating over time; A computer-implemented method according to any one of claims 1 to 5.
9. generating the enhanced STA includes removing residual noise by tracking minimum spectro-temporal energy over time; A computer-implemented method according to any one of claims 1 to 5.
10. Generating the enhanced STA includes applying a machine learning model trained with spectrotemporal amplitude data corresponding to different degrees of reverberation and other noise as input data and corresponding spectrotemporal amplitude data corresponding to only clean speech as output data. A computer-implemented method according to any one of claims 1 to 5.
11. extracting features characterizing the clean speech from the application of the machine learning model, including a low cutoff modulation frequency and a high cutoff modulation frequency; The computer-implemented method of claim 10 further comprising:
12. the calculating step includes calculating enhanced Mel-Frequency Cepstral Coefficients (MFCCs) using the enhanced STA; A computer-implemented method according to any one of claims 1 to 5.
13. the calculating step includes calculating an enhanced spectral flatness (SFT) by substituting the enhanced STA for the STA and summing values over time in calculating the SFT. A computer-implemented method according to any one of claims 1 to 5.
14. The one or more features include a spectral peak based on the sum of the power ratios of the peak band to other bands, a spectral peak based on the power ratio of the peak to the average (without the peak band), a variance or standard deviation of adjacent spectral band power, a sum or maximum of the spectral band power difference between adjacent frequency bands, a spectral variance around a spectral center, and a spectral entropy. A computer-implemented method according to any one of claims 1 to 5.
15. the determining step includes applying a machine learning model trained using one or more features of spectro-temporal amplitude data corresponding to clean speech and spectro-temporal amplitude data corresponding to different degrees of reverberation and other noise as input data, and the corresponding speech degrees as output data. A computer-implemented method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speech recognition device and signal recognition device
JP1994295196A
Sound source enhancement device, sound source enhancement learning device, sound source enhancement method and program
JP2019090930A
Sound Processing Device and Program
US20090192788A1
Sound source-separating device and sound source -separating method
US20160064000A1
Speech processing apparatus, speech processing method, and computer program product
US20160217809A1