Wind noise mitigation and speech enhancement (WMSE) for speech communications

A two-level sub-band modeling approach using specformers effectively mitigates wind noise in speech communications, improving speech intelligibility and quality by accurately separating wind noise from speech signals.

WO2026049768A1PCT designated stage Publication Date: 2026-03-05GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/056849
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-29
Filing Date
2024-11-21
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing audio processing systems struggle to effectively mitigate wind noise in speech communications, leading to compromised speech intelligibility and quality due to spectral overlap and inaccuracies in noise estimation, which results in residual errors and artifacts.

Method used

A two-level, non-linear sub-band modeling approach using specformers with non-linear and global attention mechanisms to estimate wind noise and enhance speech intelligibility, comprising a specforming wind noise modeler and a specforming speech enhancer.

Benefits of technology

The proposed method significantly improves speech intelligibility and quality by accurately separating wind noise from speech signals, reducing artifacts, and enhancing the clarity of speech in windy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056849_05032026_PF_FP_ABST
    Figure US2024056849_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques are described herein for wind noise mitigation and speech enhancement (WMSE) based on a two-level, non-linear sub-band modeling approach. The first-level model is trained to estimate a wind noise signal, W(n), from a combined speech and wind noise input, Sn(n). Directly and only removing W(n) from Sn(n) tends to yield a decent, but relatively low intelligibility speech signal. The second-level model is trained to enhance the intelligibility of the speech signal to generate a highly intelligible output speech signal S'(n). Each of the first-level and the second-level models includes a novel architecture referred to herein as a specformer, which uses a combination of non-linear and global attention mechanisms to perform the wind noise modeling and speech enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.094021-1463553 WIND NOISE MITIGATION AND SPEECH ENHANCEMENT (WMSE) FOR SPEECH COMMUNICATIONS CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202441065279, filed on August 29, 2024, and titled “WIND NOISE MITIGATION AN SPEECH ENHANCEMENT (WMSE) FOR SPEECH COMMUNICATIONS,” the content of which is herein incorporated by reference in its entirety for all purposes. BACKGROUND

[0002] People often engage in speech-related communications (e.g., phone calls, teleconferences, videoconferences, dictation, etc.) in areas prone to wind noise. For example, a person may have a phone conversation while walking outdoors, riding on or in an open vehicle, etc. Naturalness in speech tends to be compromised when communicating in windy conditions or during audio and video recording in windy environments. This loss of naturalness occurs because audio from undesired wind noise can overpower audio from desired speech at least because it typically carries high energy in the lower end of the audible spectrum, which is highly perceptible to the human auditory system. Consequently, a person may try to speak as loudly as possible to overcome the wind noise, which can create confusion in information delivery and / or other undesirable effects, a phenomenon referred to as the Lombard effect. Often, the result is the speech production process being influenced by the high-energy wind, ultimately leading to poor speech quality. SUMMARY

[0003] Systems and methods are described herein for wind noise mitigation and speech enhancement (WMSE) based on a two-level, non-linear sub-band modeling approach. The first- level model is trained to estimate a wind noise signal, W(n), from a combined speech and wind noise input, Sn(n). Directly and only removing W(n) from Sn(n) tends to yield a decent, but relatively low intelligibility speech signal. The second-level model is trained to enhance the intelligibility of the speech signal to generate a highly intelligible output speech signal S’(n). Each of the first-level and the second-level models includes a novel architecture referred to herein as a specformer, which uses a combination of non-linear and global attention mechanisms to perform the wind noise modeling and speech enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components orfeatures may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.

[0005] FIG.1 shows a plot 100 of illustrative spectral responses for typical wind and speech signals.

[0006] FIGS.2A and 2B show block diagrams of two wind noise mitigation and speech enhancement (WMSE) systems, according to embodiments described herein.

[0007] FIG.3 shows a block diagram of an illustrative specforming wind noise modeler (SWNM), according to some embodiments described herein.

[0008] FIG.4 shows a block diagram of an illustrative wind noise (WN) trained specformer including J specformer blocks.

[0009] FIG.5 shows an example implementation of a portion of an illustrative WN-trained specformer.

[0010] FIG.6 shows a block diagram of an embodiment of an adaptive filter, according to embodiments described herein.

[0011] FIG.7 shows a block diagram for an illustrative minimum phase response estimator, according to some embodiments described herein.

[0012] FIG.8 shows a block diagram of an illustrative specforming speech enhancer (SSE), for use with some embodiments described herein, such as those illustrated by FIG.2B.

[0013] FIG.9 shows a block diagram of an illustrative specforming speech enhancer (SSE), for use with some embodiments described herein, such as those illustrated by FIG.2A.

[0014] FIG.10 shows a block diagram of an illustrative partial implementation of a speech (S) trained specformer.

[0015] FIG.11 shows an illustrative training environment for training a SWNM, according to embodiments herein.

[0016] FIG.12 shows an illustrative training environment for training an SSE, such as the SSE of FIG.2A, 9, and 10, according to embodiments herein.

[0017] FIG.13 shows an illustrative training environment for training an SSE, such as the SSE of FIG.2B and 8, according to embodiments herein.

[0018] FIG.14 shows a flow diagram of an illustrative method for wind noise mitigation and speech enhancement (WMSE), according to various embodiments.

[0019] FIG.15 provides a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments. DETAILED DESCRIPTION

[0020] People often engage in speech-related communications (e.g., phone calls, teleconferences, videoconferences, dictation, etc.) in areas prone to wind noise. For example, a person may have a phone conversation while walking outdoors, riding on or in an open vehicle, etc. Naturalness in speech tends to be compromised when communicating in windy conditions or during audio and video recording in windy environments. This loss of naturalness occurs because audio from undesired wind noise can overpower audio from desired speech at least because it typically carries high energy in the lower end of the audible spectrum, which is highly perceptible to the human auditory system. Consequently, a person may try to speak as loudly as possible to overcome the wind noise, which can create confusion in information delivery and / or other undesirable effects, a phenomenon referred to as the Lombard effect. Often, the result is the speech production process being influenced by the high-energy wind, ultimately leading to poor speech quality.

[0021] FIG.1 shows a plot 100 of illustrative spectral responses for typical wind and speech signals. The plot 100 shows amplitude (in decibels) versus frequency (in kilohertz), ranging from tens of hertz (Hz) to around 25,000 Hz (25 kHz). As shown, the speech signal 110 is easily distinguishable from the wind noise signal 120 for frequencies above around 2 kHz. Generally, speech signals tend to have more energies in the range of at least around 400Hz to 1500Hz. For these lower frequencies, (e.g., in the range indicated by the dashed oval 115), there is significant overlap between the speech signal 110 and the wind noise signal 120.

[0022] This overlap can interfere with speech intelligibility during communications and / or recording in the presence of wind noise. Other practical factors can further contribute to the problem of wind noise interfering with speech intelligibility. As one example, a person may be engaged in a phone conversation using earbuds. Due to a generally low frequency response over the channel between the user’s mouth and the earbud's speaker, the dominant region of speech picked up by the earbud’s speaker tends to be highly overlapped by wind noise.

[0023] Some modern audio processing systems introduce a wind noise suppression block at the front-end to enhance the quality of the captured audio, particularly in environments prone to wind interference. This block typically includes (or is coupled with) one or more microphones and advanced digital signal processing (DSP) algorithms designed to detect and mitigate the impact of wind noise before further audio processing stages. The architecture often includes multi-channel audio capture, wind noise detection mechanisms, and adaptive filtering techniques. The DSP algorithms utilize machine learning models, such as artificial neural networks (ANNs), trained to distinguish between wind noise and speech components. The goal of such a block can be to preserve the integrity of the speech signal while minimizing the distortion introduced by the noise suppression process.

[0024] In one conventional approach, the wind noise suppression block directly estimates the clean speech from the mixed audio signal containing both speech and wind noise. Such an approach relies on machine learning models (e.g., artificial neural networks, ANNs) trained to map the noisy input to a clean speech output by learning the characteristics of both speech and wind noise. However, this approach is conventionally frustrated by the spectral overlap between speech and wind noise, particularly in the lower frequency ranges, as noted above. Since a substantial portion of the speech signal is relatively clean across most frequencies (i.e., there may be negligible overlap between the speech signal 110 and the wind noise signal 120 in most of the audible frequency range), the neural network's training error will tend to appear marginal overall, even though error remains significant in the lower frequencies. Because much of the information content in the speech signal 110 tends to be in the lower frequency range, they can be important for intelligibility and naturalness of speech. The inability to effectively reduce noise in these critical bands results in poor subjective quality, as the perceived clarity and fidelity of the speech output are compromised.

[0025] In another conventional approach, the wind noise suppression block estimates the wind noise for each frequency bin and subsequently subtracts it from the mixed signal to retrieve the clean speech. This approach typically involves spectral analysis where the energy levels of wind noise and speech are estimated separately, and the estimated noise is subtracted from the noisy signal. However, as illustrated in FIG.1, both the speech signal 110 and the wind noise signal 120 typically have high energy levels in lower frequencies. Subtracting two high-energy signals can lead to substantial residual errors due to inaccuracies in noise estimation, resulting in artifacts and distortion. Consequently, while this approach may reduce the presence of wind noise, it tends to do so at the cost of degrading the quality of the speech signal, particularly in the low-frequencyrange where human hearing is most sensitive to distortions and artifacts, thus leading to unsatisfactory subjective audio quality.

[0026] Embodiments herein describe wind noise mitigation and speech enhancement (WMSE) systems based on a two-level, non-linear sub-band modeling approach. The first-level model is trained to estimate a wind noise signal, W(n), from a combined speech and wind noise input, Sn(n). Directly and only removing W(n) from Sn(n) tends to yield a decent, but relatively low intelligibility speech signal. The second-level model is trained to enhance the intelligibility of the speech signal to generate a highly intelligible output speech signal S’(n). Each of the first-level and the second-level models includes a novel architecture referred to herein as a specformer, which uses a combination of non-linear and global attention mechanisms to perform the wind noise modeling and speech enhancement.

[0027] FIGS.2A and 2B show block diagrams of two wind noise mitigation and speech enhancement (WMSE) systems 200, according to embodiments described herein. In both WMSE systems 200, a microphone 210 is shown for reference as receiving an input signal 215. The input signal 215 is a noisy speech signal, Sn(n), that includes a wind noise component, W(n), and a speech component, S(n) (i.e., Sn(n) = W(n) + S(n)). The illustrated microphone 215 may or may not be part of the WMSE system 200, and the microphone 210 can include any suitable one or more microphones (e.g., a single microphone, a microphone array, etc.) for capturing the input signal 215.

[0028] Both WMSE systems 200 include a specforming wind noise modeler (SWNM) 220 and a specforming speech enhancer (SSE) 240. The WMSE system 200b of FIG.2B additionally includes an adaptive filter 230. The SWNM 220 includes a machine learning model (e.g., an ANN) trained to take the input signal 215 as its input and to output an estimate of W(n). For simplicity, the output of the SWNM 220 is shown nominally as the wind noise component, W(n), even though it is technically an estimate of W(n). In the WMSE system 200a of FIG.2A, the output of the SWNM 220 and the input signal 215 are both provided as inputs to the SSE 240. In the WMSE system 200b of FIG.2B, the output of the SWNM 220 and the input signal 215 are both provided as inputs to the adaptive filter 230. The output of the adaptive filter 230 is then provided as a single input to the SSE 240.

[0029] In both WMSE system 200a and WMSE system 200b, the output of the SSE 240 is a clean speech output signal 245, S’(n). Each WMSE system 200 can provide different features. For example, WMSE system 200a can be implemented with fewer processing resources because itdoes not use the adaptive filter 230. Including the adaptive filter 230 in WMSE system 200b can consume more resources but can yield a more intelligible output signal 245.

[0030] FIG.3 shows a block diagram of an illustrative specforming wind noise modeler (SWNM) 300, according to some embodiments described herein. The SWNM 300 can be an implementation of the SWNM 220 of FIG.2. As illustrated, the SWNM 300 receives input signal 215, Sn(n), and outputs an estimated wind noise output signal 345, W(n). The SWNM 300 includes a fast Fourier transform (FFT) block 310, a non-linear grouper 320, and a wind-noise- trained (WN-trained) specformer 330. The WN-trained specformer 330 includes J specformer blocks 332, where J is a positive integer.

[0031] The input signal 215 can be a stream of time-based samples as a suitable audio sampling rate. For example, a sampling rate of 24,000 samples per second (i.e., 24 kHz) can faithfully capture a maximum frequency of around 12 kHz (i.e., the Nyquist frequency is half the sampling rate), a sampling rate of 44.1 kHz can be used for audio up to around 22 kHz, etc. The FFT block 310 is configured to convert the time-domain samples into N frequency bins (labeled ‘X1’ through ‘XN’), or input-side frequency bins (Fin bins) 315. N is a positive integer typically greater than 50. In one implementation, N is 128. In another implementation, N is 256. Typically, each Fin bin 315 corresponds to a center frequency, and the center frequencies of the Fin bins 315 are evenly spaced across the represented bandwidth. For example, the total represented bandwidth (e.g., the Nyquist frequency) is 24 kHz, and the FFT block 310 generates 256 Fin bins 315 (i.e., N = 256), so that each Fin bin 315 is spaced apart by 93.75 Hz.

[0032] The N Fin bins 315 are grouped into K bin groups 325 (labeled ‘C1’ through ‘CK’) in a non-linear fashion by the non-linear grouper 320. K is an integer greater than 1 and less than N. In some implementations, K is 4. In other implementations, K is 8. The non-linear grouper 320 is configured to generate the groups based on a characterization of the spectral distribution of the wind noise signal. In particular, at least a first of the bin groups 325 is assigned a relatively small number of the Fin bins 315 corresponding to lower frequencies, and at least a second of the bin groups 325 is assigned a relatively large number of the Fin bins 315 corresponding to higher frequencies. In some implementations, the Fin bins 315 are grouped so that each bin group 325 tends to represent a similar portion of the total spectral energy of the wind noise signal. For example, each Fin bin 315 at lower frequencies will tend to have a higher amount of energy, and each Fin bin 315 at higher frequencies will tend to have a lower amount of energy, so that a smaller grouping of lower-frequency Fin bins 315 and a higher number of higher-frequency Fin bins 315 can represent similar magnitudes of spectral energy.

[0033] The WN-trained specformer 330 takes the K non-linear bin groups 325 as inputs and outputs N output-side frequency bins (Fout bins) 335. In some implementations, the Fout bins 335 are passed to an inverse FFT (IFFT) block 340. The IFFT block 340 is configured to convert the N Fout bins 335 from the frequency domain back to the time domain. For example, in context of the WMSE system 200a of FIG.2A, the output of the SWNM 300 can be passed directly to the input of the SSE 240. In some such cases, the SSE 240 can take the N Fout bins 335 directly as inputs instead of converting to the time domain and converting back to the frequency domain. In context of the WMSE system 200b of FIG.2B, the output of the SWNM 300 is passed to the input of the adaptive filter 230, which can operate in the time domain. As such, it can be desirable first to convert the output of the SWNM 300 to the time domain (via the IFFT block 340) prior to sending the output to the adaptive filter 230.

[0034] As illustrated, embodiments of the WN-trained specformer 330 include J specformer blocks 332. In some implementations, J is 4. In other implementations, J is 8. Experimentation by the inventors has revealed that 4 specformer blocks 332 tends to produce a reliably close estimation of the wind noise model (e.g., experimentation shows a relatively large improvement in performance from three to four specformer blocks 332, but a relatively small additional improvement from four to five specformer blocks 332). In general, each specformer block 332 processes each bin group 325 separately (e.g., in parallel) as a corresponding sub-band of the overall signal with localized attention. The specformer block 332 then reprocesses the sub-bands together with global attention. The output of each specformer block 332 can be a tuned version of the same N frequency bins. For example, the magnitudes of the Fin bins 315 are adjusted in each specformer block 332 to get closer to an estimate of the wind noise component of the signal.

[0035] FIG.4 shows a block diagram of an illustrative WN-trained specformer 400 including J specformer blocks 332. The WN-trained specformer 400 can be an implementation of the WN- trained specformer 330 of FIG.3 and / or can be implemented as part of the WMSE system 200 of FIG.2A or 2B. For context, FIG.4 also shows the N Fin bins 315 grouped by the non-linear grouper 320 into K bin groups 325. Each specformer block 332 can be considered essentially as a modeling stage that takes the frequency bins (X1 - XN) grouped in the K non-linear bin groups 325 (C1 – CK) as inputs and generates updates values for the frequency bins as its output. For example, the output from each stage can be a flattened vector representing the updated frequency bin values.

[0036] In each specformer block 332, there are K long short-term memory (LSTM) blocks 410 and a dense layer 420. For example, in the first specformer block 332-1, the K LSTM blocks 410are labeled 410-11 through 410-1K; in the second specformer block 332-2, the K LSTM blocks 410 would be labeled 410-21 through 410-2K; and in the Jth specformer block 332-J, the K LSTM blocks 410 would be labeled 410-J1 through 410-JK. Each LSTM block 410 is a type of recurrent neural network (RNN) designed to model sequences and capture long-term dependencies in data by using memory cells and gates. Each LSTM block 410 can be classified by three main gates: an input gate, a forget gate, and an output gate. The gates regulate the flow of information into, within, and out of the LSTM block 410, allowing the network to maintain and update a state (the memory) over time.

[0037] The input to each LSTM block 410 is a respective one of the bin group 325 inputs for that stage. For example, the input to each LSTM block 410 can be a sequence of vectors, where each vector represents a bin group 325 of frequency bins (e.g., Fin bins 315 in the first stage). As noted with respect to FIG.3, the Fin bins 315 can be derived from a FFT block, such as using a Short-Time Fourier Transform (STFT). As such, each Fin bin 315 corresponds to a specific frequency range, and its value represents the magnitude of the signal’s strength in that range at a given time step. As each LSTM block 410 processes the sequence of frequency bins, it uses its gates to selectively remember or forget information, thus learning patterns and dependencies across different time steps.

[0038] In one implementation, the output of each LSTM block 410 is a sequence of vectors, such as a sequence of predicted values or features generated from the input frequency bins. In another implementation, the output of each LSTM block 410 is a single vector, such as a predicted label or set of continuous values based on the entire input sequence. In another implementation, the output of each LSTM block 410 is a contextually relevant sequence generated based on the local attention mechanisms of the LSTM block 410. In any of those implementations, the output of the LSTM block 410 is essentially a transformation of the input sequence of frequency bins that reflects learned temporal dependencies and patterns in the data.

[0039] In each specformer block 332 (i.e., in each model stage), the outputs from the LSTM blocks 410 are passed to the dense layer 420 for that stage. The dense layer 420, sometimes referred to as a “fully connected layer,” is configured so that neuron is connected to every neuron in the previous layer. The input to the dense layer 420 can be a vector of values, each representing a feature or activation from the LSTM blocks 410. When an input vector is fed into the dense layer 420, each neuron in the layer performs a weighted sum of all the input values, adds a bias term, and passes the result through an activation function. The activation function introduces non- linearity into the model. The weights and biases are parameters that are learned during the trainingprocess through backpropagation, where the network adjusts these parameters to minimize a loss function.

[0040] The output of the dense layer 420 is another vector of values, with each value representing the activation of a neuron in the dense layer 420. The output can be configured to represent an updated set of frequency bin values. In effect, the dense layer 420 combines features learned from previous layers to form more abstract representations, thereby revealing complex patterns in the data. In essence, the dense layer transforms the input vector into an output vector through a series of learned linear transformations followed by non-linear activations.

[0041] For example, the input signal is split (e.g., by STFT, or another suitable approach), into the K bin groups 325 (sub-bands) of frequency bins. In the first specformer block 332-1, the bin groups 325 (C11 – C1K) are each fed into a respective first-stage LSTM block 410-1 (410-11 – 410-1K). Each LSTM block 410-1 processes the input sequence of its respective sub-band over time, learning temporal dependencies and patterns within the sub-band. Each LSTM block 410-1 input is a sequence of vectors, each vector corresponding to the amplitude values of the frequency bins in the sub-band at a specific time step. As the LSTM block 410-1 processes the sequence, it updates its internal state and generates a sequence of hidden states, capturing the temporal features of the input sub-band. The hidden state sequences from all the parallel LSTM blocks 410-1 are then concatenated or otherwise combined to form a feature vector, which is fed into the first stage’s dense layer 420-1. The dense layer 420-1 performs a weighted sum of its inputs, followed by a non-linear activation function, to produce an output. This output represents an integrated analysis of the temporal features extracted by the LSTM blocks 410-1 from the different frequency sub-bands.

[0042] In the next (second) specformer block 332-2, the output of the dense layer 420-1 from the first specformer block 332-1 serves as the input to second-stage LSTM blocks 410-2. For example, the LSTM blocks 410-2 receive the dense layer 420-1 output as a sequence of vectors, where each vector now represents a more abstracted and integrated feature set derived from the entire frequency spectrum. Each LSTM block 410-1 in this second specformer block 332-2 processes its input sequence, capturing higher-order temporal dependencies and patterns that span across the previously combined sub-band features. The LSTM blocks 410-2 update their internal states and output new sequences of hidden states, reflecting these higher-level temporal features. As in the first specformer block 332-1, hidden state sequences are concatenated or otherwise combined to form another feature vector, which is fed into a second-stage dense layer 420-2. Thedense layer 420-2 again performs a weighted sum of its inputs followed by a non-linear activation function to produce the second-stage output.

[0043] Additional stages of specformer blocks 332 can continue in a similar fashion up to J specformer blocks 332. The output of the Jth specformer block 332-J can be the Fout bins 335, or it can be passed to another block to generate the Fout bins 335. Between each specformer block 332, FIG.4 shows the non-linear grouper 320 in dashed lines. This is intended to represent that the output from each dense layer 420 in each of the first through (J-1)th stage is grouped into the non-linear sub-band groupings (i.e., the bin groups 325) for inputting to the corresponding LSTM blocks 410 of the subsequent specformer block 332.

[0044] FIG.5 shows an example implementation of a portion of an illustrative WN-trained specformer 500. In the illustrative implementation, the input signal 215 is a 24 kHz signal (e.g., sampled at 48 kHz). The FFT block 310 uses STFT to convert the input signal 215 into 128 Fin bins 315, so that each is evenly spaced apart by 187.5 Hz. The non-linear grouper 320 in the illustrated configuration generates four bin groups 325 (i.e., K = 4). A first bin group 325 (C1) includes the 4 lowest-frequency bins (0 Hz – 750 Hz). A next bin group 325 (C2) includes the next 20 frequency bins (750 Hz – 4.5 kHz). A next bin group 325 (C3) includes the next 24 frequency bins (4.5 kHz – 9 kHz). A final bin group 325 (C4) includes the remaining 80 highest- frequency bins (9 kHz – 24 kHz). As described above, the frequency bins are non-linearly distributed over the bin groups 325 so that a much smaller number of low-frequency bins is grouped together in the first bin group 320 relative to the much larger number of high-frequency bins grouped together in the last bin group 320. As described with respect to FIG.4, each bin group 325 is processed by a respective LSTM block 410, the outputs of the LSTM blocks 410 are combined and sent to the dense layer 420, and the dense layer 420 produces an output for the specformer block 332.

[0045] Returning to the WMSE system 200 of FIG.2B, embodiments pass the output of the SWNM 220 and the input signal 215 as inputs to the adaptive filter 230. FIG.6 shows a block diagram of an embodiment of an adaptive filter 600, according to embodiments described herein. The adaptive filter 600 can be an implementation of the adaptive filter 230 of FIG.2B. Theadaptive filter 600 is configured to have a response of ^^^^(^^^^) = ^^^^(^^^^) / ^^^^^(^^^^). Recognizing that^^^^(^^^^) = ^^^^^(^^^^) − ^^^^(^^^^), this can be further written as follows:^^^^(^^^^) = (^^^^^(^^^^) − ^^^^(^^^^)) / (^^^^^(^^^^) ) = 1 − ^^^^(^^^^) / ^^^^^(^^^^)

[0046] Taking the logarithm of both-sides yields the following:^^^^^^^^^^^^(^^^^(^^^^)) = ^^^^^^^^^^^^(^^^^^(^^^^) − ^^^^(^^^^)) − log(^^^^^(^^^^))

[0047] This can be further simplified using a cepstrum-to-impulse response transformation function ϒ as follows: ^^^^(^^^^) = ϒ(^^^^₁ − ^^^^₂),where ^^^^₁ and ^^^^₂ are cepstral coefficients corresponding to (^^^^^(^^^^) − ^^^^(^^^^)), and ^^^^^(^^^^),respectively.

[0048] Embodiments of the adaptive filter 600 response can be decomposed into minimum- phase and all-pass responses. The minimum-phase response can be estimated using cepstral coefficients (as described above) and can be applied to the noise speech before sending it to the speech enhancer for enhancement. In such an implementation, the all-pass response is not applied by the adaptive filter 600; rather it is handled by the SSE 240.

[0049] For smooth estimation, the input power spectrum of noisy speech and wind noise can be filtered using a respective first-order low-pass filter. As illustrated, each of the input signal 215 (i.e., the microphone-received speech signal 110 plus wind noise signal 120) and the wind noise signal 120 as output by the SWNM 300 are input to respective periodogram blocks 610. Each periodogram block 610 is a time-series signal processing block that estimates the power spectral density (PSD) of its respective signal. For example, each periodogram block 610 uses a discrete Fourier transform (DFT) of the time series data of it respective signal, squares the magnitude of the resulting complex numbers, and normalizes by the length of the time series to convert the time- domain sequence data into frequency-domain data to identify dominant frequencies and periodicities within the signal. Thus, a first periodogram block 610-S takes the input signal 215 as an input and outputs an estimated PSD as S’n(ω), and a second periodogram block 610-W takes the wind noise signal 120 (from the SWNM 300 output) as an input and outputs an estimated PSD as W’(ω).

[0050] Both periodogram block 610 outputs are passed to a filter response estimator 620. The filter response estimator 620 is configured to determine characteristics of a filter which, when applied to the input signal 215, would produce a desired output signal based on both of the periodogram block 610 outputs. In this case, the desired output signal is essentially the input signal 215 without the wind noise signal 120, which would be the speech signal 110. Embodiments can perform this estimation based on the spectral densities of each periodogram block 610 output (i.e., which may be the periodogram block 610 outputs themselves, also referred to as auto-spectral density, ASD) and / or the correlation between the spectral densities (i.e., cross-spectral density, CSD). For CSD, the product of the Fourier transform of one of the signals can be taken with the complex conjugate of the Fourier transform of the other signal. Based on the CSD and ASD, the filter response estimator 620 can derive the frequency response function that would modify the amplitude and phase of each frequency component of the input signal 215 to produce the speech signal 110.

[0051] In some embodiments, the adaptive filter 600 includes a regularizer block 630 after the filter response estimator 620 to impose additional constraints on the model. This can provide several features, such as reducing overfitting, enhancing generalization, improving stability, etc. Some embodiments of the regularizer block 630 can be configured to implement a forgetting factor of α, where α is a real number between 0 and 1 (or represented as a percentage from 0 to 100 percent). In each temporal frame m, a new coefficient Cm+1is computed based on the αCm-1+ (α- 1)Cm, where Cmis the coefficient in the current temporal frame and Cm-1is the coefficient for previous temporal frame. In some embodiments, α = 0.95. As such, in each temporal frame, the next coefficient is computed based on 95 percent of the prior coefficient and 5 percent of the current coefficient.

[0052] In some implementations, the regularizer block 630 performs so-called “L1 regularization,” or “lasso regression,” to add a penalty proportional to the absolute value of the coefficients, thereby leading to sparse solutions where some coefficients are exactly zero. In some implementations, the regularizer block 630 performs so-called “L2 regularization,” or “ridge regression,” which adds a penalty proportional to the square of the coefficients’ magnitudes, thereby tending to uniformly shrink coefficients and reduce the impact of noise. In some implementations, the regularizer block 630 performs a combination of L1 and L2 regularization, such as so-called “elastic net regularization.” Other implementations can perform other types of regularization, such as Tikhonov regularization.

[0053] The output of the regularizer block 630 is a regularized estimated filter response for the present temporal frame, Hm(ω). The Hm(ω) can be used to adaptively define the characteristics of a filter block 640. The filter block 640 takes each frame of the input signal 215 as its input and outputs a corresponding frame of S’(ω), which is an estimate of the speech signal 110 portion of the input signal 215 based on the estimated frame-adaptive response, Hm(ω). In context of the WMSE system 200b of FIG.2B, S’(ω) is considered as a “good enough” speech signal to be enhanced by the SSE 240. In such embodiments, the SSE 240 is effectively trained to take S’(ω) as its input and to output an enhanced speech signal that more faithfully represents speech signal110. For example, the output of the filter block 640 is a noisy or otherwise degraded version of the speech signal 110, but generally free of wind noise signal 120 artifacts.

[0054] FIG.7 shows a block diagram for an illustrative minimum phase response estimator 700, according to some embodiments described herein. The minimum phase response estimator 700 can be an implementation of the filter response estimator 620 of the adaptive filter 600 of FIG.6. The minimum phase response estimator 700 is configured based on finding a filter response H(ω) that effectively filters the wind noise portion of a noisy speech signal to generate a clean speechsignal. This can be expressed as ^^^^(^^^^) = ^^^^′^^^^(^^^^)^^^^(^^^^), which can be rewritten as follows:

[0055] As illustrated, a spectral estimate of the noisy speech signal S’n(ω) (e.g., as output from periodogram 610-S of FIG.6), and a spectral estimate of the wind noise W’(ω) (e.g., as output from periodogram 610-W of FIG.6) are received at the input of minimum phase response estimator 700. A first summer 705 is used to subtract W’(ω) from S’n(ω). A first logarithmic operator block 710-S is applied to a magnitude of S’n(ω) (|S’n(ω)|), and a second logarithmic operator block 710-W is applied to a magnitude of the output of the first summer 705 (|W’(ω)|).

[0056] A second summer 715 subtracts the output of the first logarithmic operator block 710-S from the output of the first summer 705. The combined actions of the logarithmic operator blocks 710 and the second summer 715 effectively takes the log of both sides of Equation (1), as follows: ′′ log�^^^^(^^^^)� = log(^^^^ ^^^^(^^^^) − ^^^^′(^^^^)) − log�^^^^ ^^^^(^^^^)� (2)

[0057] Accounting for using the magnitudes (absolute values), the result of Equation (2) can be expressed as HR’(ω). The implementation of FIG.7 takes advantage of cepstra of the signals. A cepstrum is a representation of a signal obtained by taking the inverse Fourier transform of the logarithm of the Fourier transform. Such an approach facilitates identification and analysis of periodicity in the signals. A “real” cepstrum focuses on the magnitude (i.e., the absolute value) of the Fourier transform, while a “complex” cepstrum includes both magnitude and phase -1information. For example, applying a Fourier transform (Ƒ) of an inverse Fourier transform (Ƒ ) to the right side of Equation (2) can be expressed as follows: log�^^^^(^^^^)� = Ƒ�Ƒ−1�log(^^^^′^^^^(^^^^) − ^^^^′(^^^^)) − log(^^^^′ ^^^^(^^^^)) (3)

[0058] Letting ^^^^ = log(^^^^′ (^^^^) − ^^^^′(^^^^)), and ^^^^ = log(^^^^′ ^^^^ ^^^^(^^^^)), Equation (3) can be furtherrewritten as log(^^^^(^^^^)) = Ƒ�^^^^ (^^^^) − ^^^^ where k1 and k are the complex cepstrumcoefficients of A and B, respectively, and assuming that the signals have both magnitude and phase information.

[0059] A relationship between the complex cepstrum coefficients k(q) and corresponding realcepstrum coefficients c(q) can be defined as k(q) = 2c(q)u(q) − u(q)^^^^(^^^^). For example, thissuggests that the complex spectrum can be derived by scaling the real cepstrum by a factor of 2, modifying it by a function (u(q)), and subtracting out the influence of u(q) at some particular frequency index controlled by the Dirac function (δ(q)) (e.g., at q = 0). The function u(q) can be a weighting or window function. This relationship between the complex and real cepstra can be used to define the real cepstrum coefficients (c1and c2), with q representing an index of the frequency domain, as follows:

[0060] Equation (4) can be simplified as follows:

[0061] A frequency-invariant constraint operator can be defined as ℓ^^^^ = [2^^^^(^^^^)^^^^(^^^^)].Replacing the relevant portion of Equation (5) with this operator yields the following:

[0062] The above relationships can be implemented using circuit blocks as illustrated in FIG.7. As noted above, the output of the second summer 715 is essentially a difference between the logs of the magnitudes of the input signals (H’R(ω)). Applying a first inverse Fourier transform block 720 to H’R(ω) effectively yields the real cepstral coefficients at each frequency index (c[q]). A multiplier 725 is used to apply the frequency-invariant constraint operator, and Fourier transform block 730 is applied to the result. This effectively implements Equation (6), yielding complex cepstral coefficients of the signal at each frequency index (Km(q)). From Equation (6), it can be seen that Km(q) represents log(H(ω)). An exponential function block 740 can be applied to Km(q) to effectively remove the log and produce H(ω), as follows:

[0063] Optionally, a second inverse Fourier transform block 750 is applied to H(ω) to yield a time-domain (or sample-domain) representation as:

[0064] Returning to FIGS.2A and 2B, embodiments of the WMSE system 200 include a specforming speech enhancer (SSE) 240. In FIG.2A, the SSE 240 receives two inputs: the input signal 215 (a noisy speech signal, Sn(n), as received by a microphone 210) and the output from the SWNM 220 (the estimated wind noise signal component, W(n), of the input signal 215). In FIG. 2B, the SSE 240 receives a single input: a signal corresponding to the sequence of output frames from the adaptive filter 230, S’’(n).

[0065] For FIG.2B embodiments, the SSE 240 can be implemented in substantially the same manner as the SWNM 220. FIG.8 shows a block diagram of an illustrative specforming speech enhancer (SSE) 800, for use with some embodiments described herein, such as those illustrated by FIG.2B. As such, the SSE 800 can be an implementation of the SSE 240 of FIG.2B. As illustrated, the SSE 800 receives S’’(n) (from the adaptive filter 230) as its input signal 815, and outputs an estimated speech component signal, which is the output signal 245, S’(n), of FIG.2B. The SSE 800 includes a fast Fourier transform (FFT) block 810, a non-linear grouper 820, and a speech-trained (S-trained) specformer 830. The S-trained specformer 830 includes G specformer blocks 832, where G is a positive integer.

[0066] In general, the architecture and operation of the SSE 800 can be similar to the architecture and operation of the SWNM embodiments described in FIGS.3 and 4. The FFT block 810 is configured to convert the time-domain input signal 815 into N frequency bins (labeled ‘X1’ through ‘XN’), or input-side frequency bins (Fin bins) 815. In some embodiments, N is the same value for both the SWNM and the SSE 800. In other embodiments, N can be a different value. The N Fin bins 815 are grouped into K bin groups 825 (labeled ‘C1’ through ‘CK’) in a non-linear fashion by the non-linear grouper 820. In some embodiments, K is the same value for both the SWNM and the SSE 800. In other embodiments, K can be a different value.

[0067] Similar to the SWNM, the non-linear grouper 820 is configured to generate the groups based on a characterization of the spectral distribution of the input signal 815. Because the W(n) signal and the S’’(n) signals may have different characteristic spectral distributions, the non-linear groupings may differ between embodiments of the SWNM and the SSE 800. For example, the SSE 800 may group the N Fin bins 815 into more or fewer bin groups 825, and any of the bin groups 825 may include more or fewer of the Fin bins 815.

[0068] The S-trained specformer 830 takes the K non-linear bin groups 825 as inputs and outputs N output-side frequency bins (Fout bins) 835. In some implementations, the Fout bins 835 are passed to an inverse FFT (IFFT) block 840. The IFFT block 840 is configured to convert the N Fout bins 835 from the frequency domain back to the time domain. As illustrated, embodimentsof the S-trained specformer 830 include G specformer blocks 832. Because the SSE 800 in this context is only has to enhance what is already an estimated speech signal (from the adaptive filter 230), the model can have a relatively low complexity. Experimentation by the inventors has revealed that as few as 2 specformer blocks 832 (i.e., G = 2) tends to produce a highly intelligible speech output signal. In another implementation, G is 3. In another implementation, G is 4.

[0069] Each specformer block 832 of the SSE 800 is configured in substantially the same manner as those of the SWNM, in accordance with the number of bin groups 825. In particular, as illustrated in FIG.4, each of K LSTMs can be configured to process a corresponding one of the K bin groups 825 separately (e.g., in parallel) as a corresponding sub-band of the overall signal with localized attention. The specformer block 832 then reprocesses the sub-bands together with global attention using a dense layer. The output of each specformer block 832 can be a tuned version of the same N frequency bins. For example, the magnitudes of the Fin bins 815 are adjusted in each specformer block 832 to get closer to an estimate of the speech component of Sn(n) (i.e., input signal 215), which is an enhanced, refined, or otherwise improved version of input signal 815.

[0070] As noted above, in context of FIG.2A embodiments, the SSE 240 receives two inputs: the input signal 215 (a noisy speech signal, Sn(n), as received by a microphone 210) and the output from the SWNM 220 (the estimated wind noise signal component, W(n), of the input signal 215). FIG.9 shows a block diagram of an illustrative specforming speech enhancer (SSE) 900, for use with some embodiments described herein, such as those illustrated by FIG.2A. As such, the SSE 900 can be an implementation of the SSE 240 of FIG.2A. As illustrated, the SSE 900 receives Sn(n) (input signal 215 of FIG.2A) and W(n) (the estimated wind noise signal output by SWNM 220) as its inputs, and outputs an estimated speech component signal, which is the output signal 245, S’(n), of FIG.2B.

[0071] The SSE 900 includes at least a fast Fourier transform (FFT) block 910-S configured to convert the time-domain noisy speech input signal Sn(n) into N frequency bins (labeled ‘Xs1’ through ‘XsN’), or speech-related input-side frequency bins (S-Fin bins) 915-S. In some embodiments, the SWNM 220 is configured to output a time-domain signal W(n). For example, referring to FIG.3, there may be an IFFT block 340 at the output side of the SWNM 300. In such embodiments, the SSE 900 also includes a second fast Fourier transform (FFT) block 910-W configured to convert a time-domain wind input signal W(n) into N frequency bins (labeled ‘Xw1’ through ‘XwN’), or wind-related input-side frequency bins (W-Fin bins) 915-W. In other embodiments, the SWNM 220 is configured to generate a flattened frequency-domain vector output. For example, referring to FIG.3, there may not be an IFFT block 340 at the output side ofthe SWNM 300. In such embodiments, the SSE 900 also includes a second fast Fourier transform (FFT) block 910-W configured to convert a time-domain wind input signal W(n) into N frequency bins (labeled ‘Xw1’ through ‘XwN’), or wind-related input-side frequency bins (W-Fin bins) 915- W. Either way, a second set of N frequency bins (labeled ‘Xw1’ through ‘XwN’), or wind-related input-side frequency bins (W-Fin bins) 915-W is provided at the input side of the SSE 900. In some embodiments, N is the same value for both the SWNM and the SSE 900. In other embodiments, N can be a different value.

[0072] The SSE 900 includes a non-linear grouper 920, which groups the N S-Fin bins 915-S and the N W-Fin bins 915-W each into K bin groups 925 (labeled ‘Cs1’ through ‘CsK’ and ‘Cw1’ through ‘CwK’, respectively). As described with reference to FIG.8, the groupings are non-linear (e.g., similar to those for the SWNM), but are tailored to a characterization of the spectral distribution of a typical speech signal. As such, the non-linear groupings may differ between embodiments of the SWNM and the SSE 900. For example, the SSE 900 may group the N Fin bins 915 into more or fewer bin groups 925, and any of the bin groups 925 may include more or fewer of the Fin bins 915.

[0073] The 2xK bin groups 925 are passed to a speech-trained (S-trained) specformer 930. The S-trained specformer 930 takes the K non-linear bin groups 925 as inputs and outputs N output- side frequency bins (Fout bins) 935. In some implementations, the Fout bins 935 are passed to an inverse FFT (IFFT) block 940. The IFFT block 940 is configured to convert the N Fout bins 935 from the frequency domain back to the time domain. As illustrated, embodiments of the S-trained specformer 930 include a three-dimensional (3D) specformer block 931 followed by H two- dimensional (2D) specformer blocks 932, where H is a positive integer. Because the SSE 900 in this context is seeking to predict the speech signal from a noisy speech signal and an estimated wind noise signal (as opposed to simply enhancing what is already an estimated speech signal, as in FIG.8), the model can have a relatively high complexity. Experimentation by the inventors has revealed that following the 3D specformer block 931 with 72D specformer blocks 932 (i.e., H = 7) tends to produce a highly intelligible speech output signal. In another implementation, G is 5. In another implementation, G is 6.

[0074] FIG.10 shows a block diagram of an illustrative partial implementation of an S-trained specformer 1000. The S-trained specformer 1000 can be an implementation of the S-trained specformer 930 of FIG.9. As illustrated, the S-trained specformer 1000 includes one 3D specformer block 931 followed by H 2D specformer blocks 932 (only one of the 2D specformer blocks 932 is explicitly shown). Similar to the specformer blocks described with reference to FIG.4, the 3D specformer block 931 includes K LSTM blocks 410-1 (410-11 through 410-1K) followed by a dense layer 420-1. In the 3D specformer block 931, each LSTM block 410-1 receives a 2xK set of inputs from a respective pairing of one of the speech-related bin groups 925- S and a corresponding one of the wind-related bin groups 925-W. For example, the same sub- band of Sn(n) and W(n) are passed to a same one of the K LSTM blocks 410-1. Thus, local attention is given by each LSTM block 410-1 to the corresponding sub-band for both input signals at once. The outputs of the LSTM blocks 410-1 in the 3D specformer block 931 are passed to the dense layer 420-1. As described above, the dense layer 420-1 applies global attention across the entire frequency band and outputs a flattened vector.

[0075] At the output of the 3D specformer block 931, there can be a flattened vector representing a single value for each of the frequency bins. This output vector can be re-grouped by the non-linear grouper 920, so that there is now a single set of bin groups, C21 – C2K (i.e., as opposed to 2xK bin groups, as previously). Subsequent specformer blocks are referred to in this context as 2D specformer blocks 932. Each 2D specformer blocks 932 can be implemented in substantially the same manner as those of the SWNM (or those described with reference to SSE 800 in FIG.8). For example, each of K LSTMs 410 (e.g., 410-21 through 410-2K in 2D specformer block 932-1) can be configured to process a corresponding one of the K bin groups 925 separately (e.g., in parallel) as a corresponding sub-band of the overall signal with localized attention. The specformer block 932 then reprocesses the sub-bands together with global attention using a dense layer (e.g., 420-2 in 2D specformer block 932-1). The output of each 2D specformer block 932 can be a tuned version of the same N frequency bins. For example, the magnitudes of the Fin bins 915 are adjusted in each 2D specformer block 932 to get closer to an estimate of the speech component of Sn(n) (i.e., input signal 215).

[0076] FIGS.11 – 13 show illustrative training environments for components of a WMSE system. FIG.11 shows an illustrative training environment 1100 for training a SWNM 220, according to embodiments herein. As illustrated, the training environment has a repository of training speech audio 1110 and a repository of training wind noise 1115. Each can be stored in any suitable storage, including any non-transitory computer-readable storage. The training speech audio 1110 is clean speech with no wind noise. The training speech audio 1110 can be actual recorded speech. In other implementations, the training speech audio 1110 can be simulated signals having the characteristics of human speech audio. In some implementations, the training speech audio 1110 is configured to be representative of different typical speech characteristics, such as different volumes, prosody, accents, syllables, etc. Similarly, in some embodiments, the training wind noise 1115 can include recordings of actual wind. In other embodiments, thetraining wind noise 1115 can include simulated signals having the characteristics of wind noise. In some implementations, the training wind noise 1115 is configured to be representative of different typical wind characteristics, such as wind howls, gusts, etc.

[0077] As illustrated, the training speech audio 1110 and the training wind noise 1115 can be combined (e.g., by a signal combiner) to generate a training Sn(n). Embodiments of the SWNM 220 are trained to take the training Sn(n) at its input and to generate an accurate wind noise estimate W’(n) of the training wind noise 1115 at its output. To that end, during training, the SWNM 220 iteratively outputs a candidate W’(n). This is compared to the training wind noise 1115 to generate an error signal. The error signal is fed back (e.g., back-propagated) to tune the parameters (e.g., the weights) of the WN-trained specformer 330 of the SWNM 220. The training continues until the error falls below a predetermined threshold, such that the candidate W’(n) generated at the output of the SWNM 220 is a sufficiently accurate representation of the training wind noise 1115.

[0078] FIG.12 shows an illustrative training environment 1200 for training an SSE 240, such as the SSE 240 of FIG.2A, 9, and 10, according to embodiments herein. As illustrated, the training environment has the repository of training speech audio 1110 and the repository of training wind noise 1115. In some implementations, the repositories are the same as those used to train the SWNM 220 in FIG.11. As in FIG.11, the training speech audio 1110 and the training wind noise 1115 can be combined (e.g., by a signal combiner) to generate a training Sn(n). It can be assumed that the SWNM 220 shown in training environment 1200 has already been trained (e.g., by the environment of FIG.11) to generate an accurate representation of the training wind noise 1115 responsive to receiving the training Sn(n).

[0079] Embodiments of the SSE 240 are trained to take, as inputs, the training Sn(n) and the W’(n) from the output of the trained SWNM 220, and to generate a candidate estimate of the training speech audio 1110. To that end, during training, the SSE 240 iteratively outputs a candidate S’(n). This is compared to the training speech audio 1110 to generate an error signal. The error signal is fed back (e.g., back-propagated) to tune the parameters (e.g., the weights) of the S-trained specformer 930 of the SSE 240. The training continues until the error falls below a predetermined threshold, such that the candidate S’(n) generated at the output of the SSE 240 is a sufficiently accurate representation of the training speech audio 1110.

[0080] FIG.13 shows an illustrative training environment 1300 for training an SSE 240, such as the SSE 240 of FIG.2B and 8, according to embodiments herein. As illustrated, the training environment has the repository of training speech audio 1110 and the repository of training windnoise 1115. In some implementations, the repositories are the same as those used to train the SWNM 220 in FIG.11. As in FIG.11, the training speech audio 1110 and the training wind noise 1115 can be combined (e.g., by a signal combiner) to generate a training Sn(n). It can be assumed that the SWNM 220 shown in training environment 1300 has already been trained (e.g., by the environment of FIG.11) to generate an accurate representation of the training wind noise 1115 responsive to receiving the training Sn(n).

[0081] The environment includes an adaptive filter 230. As described with reference to FIGS. 2B, 6, and 7, the adaptive filter 230 takes the training Sn(n) and the W’(n) from the output of the trained SWNM 220 as inputs and generates a “good enough” speech audio signal S’’(n). S’’(n) is generated by adaptively filtering the training Sn(n) for frame-by-frame removal of the estimated wind noise. In this training environment 1300, the SSE 240 is trained to take S’’(n) from the adaptive filter 230 as its input and to generate the training speech audio 1110 at its output. The training speech audio 1110 is essentially an enhanced (i.e., cleaned up) version of S’’(n). To that end, during training, the SSE 240 iteratively outputs a candidate S’(n) from the S’’(n). This is compared to the training speech audio 1110 to generate an error signal. The error signal is fed back (e.g., back-propagated) to tune the parameters (e.g., the weights) of the S-trained specformer 930 of the SSE 240. The training continues until the error falls below a predetermined threshold, such that the candidate S’(n) generated at the output of the SSE 240 is a sufficiently accurate representation of the training speech audio 1110.

[0082] FIG.14 shows a flow diagram of an illustrative method 1400 for wind noise mitigation and speech enhancement (WMSE), according to various embodiments. Embodiments begin at stage 1404 by receiving a noisy speech signal Sn(n) that includes both a speech audio component S(n) and a wind noise component W(n). For example, Sn(n) is received from one or more microphones of a portable (e.g., wearable) audio device).

[0083] At stage 1408, embodiments can generate an estimated wind noise signal W’(n) by a trained specforming wind noise modeler (SWNM). As described herein, the generating at stage 1408 is based on applying local attention to several non-linearly allocated first sub-bands of Sn(n) and applying global attention across the first sub-bands. For example, the generating at stage 1408 can include: converting Sn(n) to frequency bins and grouping the frequency bins non-linearly into bin groups, each corresponding to a respective one of the first sub-bands. As described herein, the non-linear grouping can be in accordance with a characteristic spectral distribution of wind noise. In some implementations, the trained SWNM includes several specformer blocks. In each specformer block, the generating in stage 1408 can include: generating, by each of multiple longshort-term memory (LSTM) blocks, a corresponding one of multiple sub-band outputs based on applying the local attention to a corresponding one of the bin groups; and generating a corresponding specformer output representing updated values for the frequency bins by applying the global attention to the sub-band outputs.

[0084] At stage 1412, embodiments can output an estimated speech signal S’(n) by a trained specforming speech enhancer (SSE) coupled with the SWNM. As described herein, the outputting at stage 1412 is based on applying local attention to each of several non-linearly allocated second sub-bands of a combination of Sn(n) and W’(n) and applying global attention across the second sub-bands. In some embodiments, the outputting at stage 1412 includes converting Sn(n) to a plurality of first frequency bins (S-bins) and grouping the S-bins non-linearly into a plurality of first bin groups (S-groups), each corresponding to a respective one of the second sub-bands. In some such embodiments, the outputting at stage 1412 further includes: generating, by a three- dimensional (3D) specformer block, initial values for combined frequency bins (C-bins) based on applying the local attention to each of the non-linearly allocated second sub-bands of the combination of Sn(n) and W’(n) and applying the global attention across the second sub-bands; and generating, in each of multiple two-dimensional (2D) specformer blocks, respective updated values for the C-bins based on applying the local attention to each of the non-linearly allocated second sub-bands as applied to the C-bins and applying the global attention across the second sub- bands.

[0085] Some embodiments of the method 1400 further include adaptively filtering Sn(n) based on a temporal frame-wise spectral evaluation of W’(n) to generate a wind-free speech estimate S’’(n). In such embodiments, the outputting at stage 1412 can include: converting S’’(n) to multiple frequency bins and grouping the frequency bins non-linearly into multiple bin groups, each corresponding to a respective one of the second sub-bands; and iteratively generating, by each of multiple specformer blocks, a corresponding specformer output based on applying the local attention to each of the bin groups to generate a respective one of multiple sub-band outputs and applying the global attention to the sub-band outputs.

[0086] FIG.15 provides a schematic illustration of an illustrative computational system 1500 that can implement various system components and / or perform various steps of methods provided by various embodiments. Embodiments of the computational system 1500 can be integrated in a portable audio device, such as an earbud, headset, audio recorder, cellphone, smartphone, etc. Embodiments of the computational system 1500 can implement some or all of the wind noise mitigation and speech enhancement (WMSE) systems described herein. Additionally oralternatively, embodiments can implement any of the training environments described herein, and / or can execute any of the machine learning models described herein. FIG.15 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG.15, therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.

[0087] The computational system 1500 is shown including hardware elements that can be electrically coupled via a bus 1505 (or may otherwise be in communication, as appropriate). The hardware elements may include one or more processors 1510, including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); one or more input devices 1515; and one or more output devices 1520.

[0088] The computational system 1500 may further include (and / or be in communication with) one or more non-transitory storage devices 1525, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory (“RAM”), and / or a read-only memory (“ROM”), which can be programmable, flash-updateable and / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like.

[0089] The computational system 1500 can also include a communications subsystem 1530, which can include, without limitation, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth^ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 1530 supports multiple communication technologies. Further, the communications subsystem 1530 can provide communications with one or more networks.

[0090] In many embodiments, the computational system 1500 will further include a working memory 1535, which can include a RAM or ROM device, as described herein. The computational system 1500 also can include software elements, shown as currently being located within the working memory 1535, including an operating system 1540, device drivers, executable libraries, and / or other code, such as one or more application programs 1545, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can beimplemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 1540 and the working memory 1535 are used in conjunction with the one or more processors 1510 to implement some or all of the wind noise mitigation and speech enhancement (WMSE) components, such as the SWNM 220, the SSE 240, and / or the adaptive filter 230.

[0091] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 1525 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 1500. In other embodiments, the storage medium can be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instructions / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 1500 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 1500 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.

[0092] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.

[0093] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 1500) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 1500 in response to processor 1510 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 1540 and / or other code, such as an application program 1545) contained in the working memory 1535. Such instructions may be read into the working memory 1535 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 1525. Merely by way of example, execution of the sequences of instructions contained in the workingmemory 1535 can cause the processor(s) 1510 to perform one or more procedures of the methods described herein.

[0094] The terms “machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 1500, various computer-readable media can be involved in providing instructions / code to processor(s) 1510 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer- readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 1525. Volatile media include, without limitation, dynamic memory, such as the working memory 1535. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.

[0095] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 1510 for execution. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 1500. The communications subsystem 1530 (and / or components thereof) generally will receive signals, and the bus 1505 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 1535, from which the processor(s) 1510 retrieves and executes the instructions. The instructions received by the working memory 1535 may optionally be stored on a non-transitory storage device 1525 either before or after execution by the processor(s) 1510.

[0096] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. For example, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.

Claims

WHAT IS CLAIMED IS:

1. A wind noise mitigation and speech enhancement (WMSE) system comprising: a specforming wind noise modeler (SWNM) comprising a wind noise-trained (WN- trained) specformer trained to output an estimated wind noise signal W’(n) based on receiving a noisy speech signal Sn(n) that includes both a speech audio component S(n) and a wind noise component W(n), such that W’(n) is an estimate of W(n); and a specforming speech enhancer (SSE) coupled with the SWNM and comprising a speech-trained (S-trained) specformer trained to output an estimated speech signal S’(n) based on the noisy speech signal Sn(n) and the W’(n).

2. The WMSE system of claim 1, wherein: the SWNM further comprises: a fast Fourier transform (FFT) block to convert Sn(n) to a plurality of frequency bins; and a non-linear grouper coupled with the FFT block to group the plurality of frequency bins non-linearly into a plurlaity of bin groups, each corresponding to a respective sub-band of Sn(n); and the WN-trained specformer is coupled with the non-linear grouper and comprises an artificial neural network (ANN) trained to generate W’(n) based on applying local attention to each of the plurality of bin groups and applying global attention across the plurality of bin groups.

3. The WMSE system of claim 2, wherein the WN-trained specformer comprises: a plurality of specformer blocks, each including a plurality of long short-term memory (LSTM) blocks and a dense layer, wherein each LSTM block is configured to receive a corresponding one of the plurality of bin groups and to generate a corresponding one of a plurality of sub-band outputs based on applying the local attention to the corresponding one of the plurality of bin groups, and wherein the dense layer is configured to receive the plurality of sub-band outputs and to generate a corresponding specformer output representing updated values for the plurality of frequency bins.

4. The WMSE system of claim 2, wherein: the plurality of frequency bins is N frequency bins;the plurlaity of bin groups is K bin groups; each of N and K is an integer greater than 1; and N / K is at least 10.

5. The WMSE system of claim 2, wherein: the WN-trained specformer is configured to generate a frequency-domain output; and the SWNM further comprises: an inverse fast Fourier transform (IFFT) block coupled with the WN-trained specformer to convert the frequency-domain output to W’(n) as a time-domain signal.

6. The WMSE system of claim 1, wherein: the SSE is configured to receive Sn(n) and to receive W’(n) from the SWNM; the SSE further comprises: a fast Fourier transform (FFT) block to convert at least Sn(n) to N first frequency bins (S-bins), N being an integer greater than 50; and a non-linear grouper to group the N S-bins non-linearly into a first K bin groups (S-groups) and to group N second frequency bins (W-bins) non-linearly into a second K bin groups (W-groups), wherein K is an integer greater than 1, the W-bins correspond to a frequency-domain representation of W’(n), and each of the K S-groups and each of the K W-groups corresponds to a respective one of K sub-bands of Sn(n); and the S-trained specformer is coupled with the non-linear grouper and comprises an artificial neural network (ANN) trained to generate S’(n) based on applying local attention to each of the K sub-bands and applying global attention across the K sub-bands.

7. The WMSE system of claim 6, wherein: the SSE is configured to receive W’(n) from the SWNM as a time-domain signal; and the FFT block is further to convert W’(n) to the N W-bins.

8. The WMSE system of claim 6, wherein the S-trained specformer comprises: a three-dimensional (3D) specformer block configured to generate K first sub-band outputs based on applying the local attention to each of K sub-band inputs, each corresponding to a respective one of the K S-bins and a respective one of the K W-bins, and to generate initial valuesfor N combined frequency bins (C-bins) by applying the global attention to the K first sub-band outputs; and H two-dimensional (2D) specformer blocks, each configured to generate K second sub-band outputs based on applying the local attention to each of K sub-band portions of the N C- bins as output by a respective preceding block and to generate updated values for the N C-bins by applying the global attention to the K second sub-band outputs, wherein the respective preceding block for a first of the H 2D specformer blocks is the 3D specformer block, and the respective preceding block for each hth of the H 2D specformer blocks being the (h-1)th 2D specformer block, h being an integer between 2 and H.

9. The WMSE system of claim 8, wherein: the 3D specformer block comprises K 2D long short-term memory (LSTM) blocks and a first dense layer, and each kth 2D LSTM block is configured to generate a kth one of the K first sub-band outputs based on applying the local attention concurrently to a corresponding kth W-group of the K W-groups and a corresponding kth S-group of the K S-groups, k being an integer from 1 to K.

10. The WMSE system of claim 1, further comprising: an adaptive filter configured to receive Sn(n) and to receive W’(n) from the SWNM and to generate a wind-free speech estimate S’’(n) based on temporal frame-wise removal of W’(n) artifacts from Sn(n).

11. The WMSE system of claim 10, wherein: the SSE further comprises: a fast Fourier transform (FFT) block to convert S’’(n) to a plurality of frequency bins; and a non-linear grouper coupled with the FFT block to group the plurality of frequency bins non-linearly into a plurality of bin groups, each corresponding to a respective sub-band of S’’(n); and the S-trained specformer is coupled with the non-linear grouper and comprises an artificial neural network (ANN) trained to generate S’(n) based on applying local attention to each of the plurality of bin groups and applying global attention across the plurality of bin groups.

12. The WMSE system of claim 11, wherein the S-trained specformer comprises:a plurality of specformer blocks, each including a plurality of long short-term memory (LSTM) blocks and a dense layer, wherein each LSTM block is configured to receive a corresponding one of the plurality of bin groups and to generate a corresponding one of a plurality of sub-band outputs based on applying the local attention to the corresponding one of the plurality of bin groups, and wherein the dense layer is configured to receive the plurality of sub-band outputs and to generate a corresponding specformer output representing updated values for the plurality of frequency bins.

13. A portable audio device comprising: one or more microphones to capture Sn(n); and the WMSE system of claim 1 coupled with the one or more microphones.

14. A method for wind noise mitigation and speech enhancement (WMSE), the method comprising: receiving a noisy speech signal Sn(n) that includes both a speech audio component S(n) and a wind noise component W(n); generating, by a trained specforming wind noise modeler (SWNM), an estimated wind noise signal W’(n) based on applying local attention to each of a plurality of non-linearly allocated first sub-bands of Sn(n) and applying global attention across the first sub-bands; and outputting, by a trained specforming speech enhancer (SSE) coupled with the SWNM, an estimated speech signal S’(n) based on applying local attention to each of a plurality of non-linearly allocated second sub-bands of a combination of Sn(n) and W’(n) and applying global attention across the second sub-bands.

15. The method of claim 14, wherein the generating W’(n) comprises: converting Sn(n) to a plurality of frequency bins; and grouping the plurality of frequency bins non-linearly into a plurality of bin groups, each corresponding to a respective one of the first sub-bands.

16. The method of claim 15, wherein: the trained SWNM comprises a plurality of specformer blocks; and in each specformer block, the generating W’(n) further comprises: generating, by each of a plurality of long short-term memory (LSTM) blocks, a corresponding one of a plurality of sub-band outputs based on applying the local attention to a corresponding one of the plurality of bin groups; andgenerating a corresponding specformer output representing updated values for the plurality of frequency bins by applying the global attention to the plurality of sub- band outputs.

17. The method of claim 14, wherein the outputting S’(n) comprises: converting Sn(n) to a plurality of first frequency bins (S-bins); and grouping the S-bins non-linearly into a plurality of first bin groups (S-groups), each corresponding to a respective one of the second sub-bands.

18. The method of claim 17, wherein the outputting S’(n) further comprises: generating, by a three-dimensional (3D) specformer block, initial values for combined frequency bins (C-bins) based on applying the local attention to each of the plurality of non-linearly allocated second sub-bands of the combination of Sn(n) and W’(n) and applying the global attention across the second sub-bands; and generating, in each of a plurality of two-dimensional (2D) specformer blocks, respective updated values for the C-bins based on applying the local attention to each of the plurality of non-linearly allocated second sub-bands as applied to the C-bins and applying the global attention across the second sub-bands.

19. The method of claim 14, further comprising: adaptively filtering Sn(n) based on a temporal frame-wise spectral evaluation of W’(n) to generate a wind-free speech estimate S’’(n), wherein the outputting S’(n) comprises: converting S’’(n) to a plurality of frequency bins; grouping the plurality of frequency bins non-linearly into a plurality of bin groups, each corresponding to a respective one of the second sub-bands; and iteratively generating, by each of a plurality of specformer blocks, a corresponding specformer output based on applying the local attention to each of the plurality of bin groups to generate a respective one of a plurality of sub-band outputs and applying the global attention to the plurality of sub-band outputs.

20. A computer program comprising instructions for implementing the method of claim 14.

Citation Information

Patent Citations

  • Low power multi-stage selectable neural network suppression

    US20220092389A1

  • ADL-UFE: all deep learning unified front-end system

    US20230154480A1

  • Neural-network-based approach for speech denoising statement regarding federally sponsored research

    US20230306981A1